EconBase
← Back to paper

Improving control over unobservables with network data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,587 characters · 13 sections · 31 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Improving control over unobservables with network data

abstractThis paper develops a method to conduct causal inference in the presence of unobserved confounders by leveraging networks with homophily, a frequently observed tendency to form edges with similar nodes. I introduce a concept of asymptotic homophily, according to which individuals' selectivity scales with the size of the potential connection pool. It contributes to the network formation literature with a model that can accommodate common empirical features such as homophily, degree heterogeneity, sparsity, and clustering, and provides a framework to obtain consistent estimators of treatment effects that are robust to selection on unobservables. I also consider an alternative setting that accommodates dense networks and show how selecting linked individuals whose observed characteristics made such a connection less likely delivers an estimator with similar properties. In an application, I recover an estimate of the effect of parental involvement on students' test scores that is greater than that of OLS, arguably due to the estimator's ability to account for unobserved ability. {\bf Keywords}: Causal Inference, Networks, Selection on unobservables, Homophily.

Introduction

Estimating the effect of a treatment is a frequent goal in economics and social sciences. A common challenge is the presence of unobserved confounders that threaten the validity of the unconfoundedness assumption, which is typically necessary to perform inference with standard methods. Tools that strengthen control over variables that affect outcomes but are hard to measure, such as ability, culture, work ethic, tastes, etc., are thus particularly valuable.

Recently, networks and datasets with a spatial structure have become increasingly available to researchers, providing new avenues for research. Homophily or assortative matching is an ubiquitous feature of empirical networks: nodes tend to associate with similar nodes lazarsfeld1954friendship, clark1992friendship, case1993budget, mcpherson2001birds, moody2001race, currarini2009economic, boucher2017my, dzemski2019empirical. As zeleneev2020identification note, homophily is likely to also operate through unobserved factors. An example is the tendency for people to form friendship ties based on ability clark1992friendship, burgess2011school, boutwell2017general, a variable typically unavailable to the researcher.

Homophily then generates opportunities to create unobservable-adjusted comparison groups. For instance, if we are interested in the effect of parental involvement on student test scores, we may be concerned about unobserved confounding from differences in student ability. However, if students of similar ability are likely to be friends clark1992friendship, burgess2011school, boutwell2017general, the omitted variable bias can be reduced by comparing connected students.

The paper develops the idea that homophilic networks can be exploited to derive consistent estimators of treatment effects in the presence of unobserved confounders. This is formally done under two main frameworks: either the network is sparse and homophily captures the essence of link formation, or the network is dense and features homophily at least in the unobservables.

In the former case, I let the the probability of link formation vary with the size of the network: people become pickier to limit the number of connections or improve their average quality as the network expands. This is consistent with the common view that the average degree should not increase proportionally to the size of the network and that most networks are sparse. As people are able to form increasingly better matches with a larger pool of potential neighbors, they become more selective because of decreasing benefits per additional match, preference for quality of matches, or limited resources to devote to additional connections. Improved match quality has been documented in a few instances, e.g., dauth2022matching analyze worker sorting and find evidence of stronger assortative matching in larger cities.

This modeling strategy accomplishes two things. First, the approach provides an asymptotic approximation that does not render the mechanism of network formation negligible in the limit: selectivity scales with the size of the connection pool, in the spirit of drifting sequences. As such, it contributes to the network formation literature with a model that can accommodate common features of empirical networks such as homophily, sparsity, clustering, etc. Second, it is sufficient to establish consistency and asymptotic normality for estimators that use $m^{\mbox{th}}$-order connections or people with more than $c$ connections in common as comparison groups.

In an alternative framework, I focus on dense networks and show that comparison groups can be derived by comparing people whose observables increasingly differ but nevertheless connect. This hinges on the following intuition: if there is no observed rationale for two people being friends, the reason for their friendship is more likely to lie in the unobserved world. If two people are connected despite their observables indicating that such a link was unlikely, they are more likely to be close in terms of unobservables. By suitably manipulating a discrepancy in observables and letting it grow with sample size, consistent estimators can be constructed.

I provide results that allow for the estimation of the Conditional Average Treatment Effect (CATE), which provides a way to describe the heterogeneity of the treatment effect for sampled individuals. The conditional average effect may be the end goal of the analysis (when a specific unit is targeted for treatment or policy) or may be a prelude to aggregation to the Average Treatment Effect (ATE).

I define a general form of CATE estimator as a function of a group of counterfactual observations to be determined, then propose different choices to deal with different empirical issues. In all cases, estimators isolate increasingly better counterfactuals as to recover the CATE asymptotically. I show that the proposed estimators of the (C)ATE are asymptotically normal, enabling statistical inference.

Finally, I demonstrate the feasibility and effectiveness of the method through both simulations and an empirical application. In the application, I obtain an estimate of the effect of parental involvement on students' test scores that suggests a greater impact than OLS does, arguably due to the estimator's ability to account for unobserved ability and motivation.

\paragraph{Related literature}

The paper is at the intersection of the literature on networks jackson2010social, graham2015methods, de2017econometrics, newman2018networks, in particular those featuring homophilic network formation boucher2015structural, graham2016homophily, graham2017econometric, demirer2019partial, gao2020nonparametric, mele2022structural, and estimation of treatment effects imbens2004nonparametric, imbens2009recent, imbens2015causal.

In a related paper, auerbach2022identification considers a partially linear outcome regression where the nonlinear term depends on an unobserved variable. Using information from a network whose formation hinges on the unobserved variable, he is able to recover consistent estimates of regression coefficients under general assumptions. See also goldsmith2013social, hsieh2016social, johnsson2021estimation, who use related frameworks and provide a way to analyze peer effects.

Through the help of a pseudo-distance, zeleneev2020identification devises a method to identify agents with similar values of latent fixed effects, which allows him to estimate parameters of interest while controlling for unobserved heterogeneity. demirer2019partial provides partial identification results in linear models under homophilic behavior and proposes a comprehensive nomenclature for homophily.

The present paper considers general outcome equations in a causal inference framework, at the expense of some generality in the network formation process. Specifically, I consider a nonparametric potential outcome setup, but I impose structure on network formation, especially homophily in the unobservables. The potential outcome framework is suitable to discuss causality issues, and the approach explicitly deals with the common concerns of treatment effect heterogeneity and nonlinearities. In addition, the method circumvents the need to define and estimate equivalent classes and focuses on the common case of sparse networks, in contrast to previous papers. Finally, homophilic structures allow for the use of higher-order neighbors or friends in common through triangular inequality relationships, which leads to a class of intuitive estimators that are easy to implement.

Improving control over unobservables using network data

Notation and assumptions

The sample is a cross-section of $n$ individuals. The treatment status of individual $i$, $T_i \in \{0, 1\}$, and the corresponding outcome, $Y_i = Y_i(T_i)$ with the potential outcome notation neyman1923application, rubin1974estimating, are observed. As the notation for the outcome suggests, the Stable Unit Treatment Value Assumption (SUTVA) is maintained throughout.

The covariates, $X = (X^o, X^u) \in \mathcal{X}^o \times \mathcal{X}^u \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \mathcal{X} \subset \mathds{R}^d$, are divided into observed variables, $X^o$, and unobserved variables, $X^u$. There is a norm $\Vert \cdot \Vert$ on $\mathds{R}^d$ (with some abuse of notation, this will be used to represent the norm on $\mathcal{X}^o$ or $\mathcal{X}^u$), typically Euclidean. I focus on continuously distributed covariates $X$, though discrete variables can be accommodated -- typically under weaker conditions since concerns such as asymptotic bias disappear. To avoid technical difficulties with vanishing denominators, it will be convenient to assume that covariates have a smooth density bounded from below. I thus make the following assumption throughout the analysis:

assumption[Existence of bounded densities] \ \\ The joint distribution of the covariates admits a density $f$ with respect to Lebesgue measure. On the compact $\mathcal{X}$, the density is continuously differentiable and satisfies $f \geq \underline{f}$ for some positive $\underline{f}$.

Draws of $(Y_i, T_i, X_i)$ are i.i.d. and realizations of a random variable are denoted by the corresponding lower-case letter. $B_r(x)$ denotes a ball of radius $r$ centered at $x$. $C$ represents a generic (positive) constant.

A network is given through a (binary) weighting/link matrix $W$, of size ($n \times n$). The neighborhood $\mathcal{N}(i)$ refers to the links, friends, or connections of the node or individual $i$, i.e $\mathcal{N}(i) \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \{j \in \{1, \ldots, n\}\vert W_{ij} = 1\}$, $\mathcal{N}_t(i)$ denotes neighbors with a specific treatment status $t$, i.e $\mathcal{N}_t(i) \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \{j \in \{1, \ldots, n\}\vert W_{ij} = 1, T_j = t\}$. These definitions extend to higher-order neighbors, say of order $m$, which are denoted by $\mathcal{N}_t^m(i)$. Connections in common are given by $\mathcal{N}_t(i;j) \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \mathcal{N}_t(i) \cap \mathcal{N}_t(j)$.

The goal is to conduct inference about treatment effects. In particular, I develop inference methods for the Conditional Average Treatment Effect (CATE), $\operatorname{CATE}(x_i) \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \mathds{E}[Y_i(1)-Y_i(0)\vert X_i=x_i]$, and then for the Average Treatment Effect (ATE), $\operatorname{ATE} \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \mathds{E}[Y_i(1)-Y_i(0)]$. Although I focus on average treatment effects, the insights can be exploited to obtain, e.g., quantiles of treatment effects or the average effect on the treated. The usual statement about omitting ‘almost surely’ qualifiers, in particular pertaining to conditional expectations, applies.

The following core assumptions are maintained throughout the paper:

assumption[Causal Inference] \ \\ a) Unconfoundedness: $(Y_i(1), Y_i(0)) \perp \!\!\! \perp T_i \vert X_i$ \\ b) Overlap: $0<C<\mathds{P}[T_i = 1\vert X_i] < 1-C < 1$

These two assumptions are ubiquitous in the treatment effect literature, although this version of unconfoundedness conditions on $X$ instead of $X^o$. It is thus only assumed that treatment is independent of potential outcomes when conditioned on individual characteristics, including unobserved ones. Since covariates that may influence selection into treatment, such as ability, work ethic, or personal preferences, are typically unobserved, this is often a valuable relaxation: selection on some unobservables is allowed.

Network formation

Let $i, j$ be two individuals and $i \neq j$. I focus on link-formation models of the type

equation[equation omitted — 123 chars of source]

where $w_n: \mathds{R}^+ \rightarrow [0;1]$ is a decreasing function that satisfies $\lim_{x \rightarrow \infty} w_n(x) = 0$. Typically, $w_n$ would decrease with $n$ to accommodate network sparsity, e.g., $w_n(x) = \max\{1-s_n x, 0\}$ or $e^{- s_n \frac{1}{2} x^2}$. The function $h$ is arbitrary and $\eta_{ij}=\eta_{ji}$ are independent uniform\footnote{Since any inverse cumulative distribution function can be applied on both sides to generate any distribution, the uniform assumption is just a normalization.} shocks, drawn independently of $(X_i, X_j, T_i, T_j, Y_i, Y_j)$. The dimensionality of the unobserved variables is arbitrary and matters only for rates of convergence.

Dyadic network formation processes are common in the literature, e.g., graham2017econometric, gao2020nonparametric, zeleneev2020identification, auerbach2022identification, johnsson2021estimation. Compared to more general specifications (for example, auerbach2022identification posits that links are formed whenever $\eta_{ij} \leq w(X_i, X_j)$ and only imposes a weak continuity assumption on $w$), the model ((ref)) adds some separability and homophily in the unobservables.

remarkThere may be another type of unobservable that reflects factors such as popularity or expansiveness. Such factors can be accommodated by replacing the left-hand side of (1) with $\tilde{w}_n\left(h(X_i^o; X_j^o) + \Vert X_i^u - X_j^u\Vert, A_i, A_j\right)$, where $A_i$ and $A_j$ are unobserved heterogeneities that represents expansiveness.\footnote{Since links are undirected, $A_i$ and $A_j$ should enter $\tilde{w}$ symmetrically.} The model can then be interpreted as extending graham2017econometric's (see also dzemski2019empirical in the case of directed links): compared to models where links are formed whenever $X_{ij}'\theta + A_i + A_j$ exceeds (logistic) errors, the model relaxes slightly the functional form and distribution of errors, and introduces unobserved components of $X$ that satisfy homophilic restrictions. The two frameworks coincide if we parametrize $h(X_i, X_j) + \Vert X_i^u - X_j^u\Vert \equiv X_{ij}' \theta$ where $X_{ij}$ is obtained from a transformation of $(X_i, X_j)$ such that the unobserved components are featured homophilically, and we let $\tilde{w}(z, a, b) \equiv \mbox{Logit}(x+a+b)$. On the other hand, the model also allows introducing heterogeneities in new ways. A variant that may be of interest specifies the argument of the link function as $\left(\frac{h(X_i^o; X_j^o) + \Vert X_i^u - X_j^u\Vert}{(A_i A_j)^{1/d}}\right)$. Here, expansiveness is featured in a multiplicative way, which may be easier to interpret as a selectivity scale. In the limit of the asymptotic homophily model to be developed in the next Section, it would correspond to a scaling factor for probabilities. A person with characteristics $(x_i, 2 a_i)$ is twice as expansive and has twice the probability of forming a link than a person with characteristics $(x_i, a_i)$ as $n \rightarrow \infty$. The analysis can be applied to such functions $\tilde{w}$ under the assumption that expansiveness does not affect the outcome of interest. The results of the next sections directly apply if those adjustments are made: \begin{enumerate} • In the asymptotic framework of Section 2.3, the dependence on $n$ is through the first argument of $\tilde{w}$; in Graham's model, this is interpreted as letting $\theta$ drift: links are formed whenever $X_{ij}' \theta_n$ is lower than $U_{ij} + A_i + A_j$, where $U$ is logistic and the unobserved heterogeneities may be correlated with $X$; • In the dense framework of Section 2.4, the function $\tilde{w}(z, a, b)$ can be lower bounded by a function $\underline{\tilde{w}}(z)$ that shares the properties required of Section 2.4's $w$; in Graham's model, this is interpreted as imposing bounds on $A_i$. \end{enumerate}

Another feature of this network formation is the explicit dependence of $w$ on network size, allowing for sparse networks. This is often the empirically relevant setup since the mean degree of a node is rarely expected to scale with the size of the network jackson2010social, newman2018networks.

The function $h$ may feature homophily as well\footnote{In this case in particular, it may make sense to consider variables whose variance has been normalized to put them on the same scale. Nevertheless, the results hold if the norms weight each dimension differently as to reflect stronger selection in some covariates.} in which case $h(X_i^o; X_j^o) = \Vert X_i^o - X_j^o\Vert$, but other forms may be more appropriate in some applications. For instance, some work relationships may warrant skill complementarity, in which case there is non-(possibly anti-) homophilic selection in a covariate.

The model can be given the usual interpretation of ‘link creation under a mutual positive utility of forming a link’ ($w-\eta$ then reflecting utility; see, e.g., jackson2010social), where people derive more utility from interacting with similar individuals, or rationalizes the idea that people with similar characteristics are more likely to meet and thus to form a connection. Nevertheless, since the network is primarily seen as information to draw from, this rationale may not be necessary. For instance, if individuals end up developing similar characteristics after randomly forming connections, a researcher that observes the network after covariates have evolved could use the present framework. In other words, ((ref)) need not be the structural equation for network formation, but should approximate the relationship between the links and the covariates relevant to selection at the time of observation.

I first explore the case of asymptotic homophily, i.e., homophilic behavior is the core mechanism of network formation and individuals' selectivity is tied to the size of their potential matching pool. This provides an asymptotic theory when homophilic behavior is pronounced relative to network size and formalizes the intuition that connections among individuals can be used to form comparison groups. It also shows the identifying power of homophilic restrictions under simpler conditions than the results of the later sections and provides a network formation model compatible with many empirically relevant features.

A second framework is discussed in Subsection (ref), possibly letting the link formation be independent of sample size and thus accommodating denser networks.

Asymptotic homophily

The asymptotic homophily framework

Suppose that associations are captured by homophily, the probability of a connection is decreasing in $\Vert X_i - X_j\Vert$, and the network is sparse: the average degree of a node increases arbitrarily slowly. As people are able to form increasingly better matches with a larger pool of potential neighbors, they become more selective because of decreasing benefits per additional match, preference for quality of matches, or limited resources to devote to additional connections.

To reflect this behavior, the sequence of functions $w_n$ must satisfy two conditions. First, the sequence must be decreasing in order to decrease the probability of forming connections as $n$ rises. Homophily further suggests that people penalize dissimilar individuals increasingly more harshly so that the average match quality (in terms of homophilic preferences) increases. This is in line with empirical evidence in some applications, such as stronger assortative matching in larger cities dauth2022matching.

Functions of the form $w_n(x) \geq g(s_n x)$ are consistent with such behavior -- homophily becomes more prevalent as $n$ rises -- irrespective of the exact form of $w_n$ (or $g$). I adopt the following definitions:

definition[Asymptotic homophily] a) Network formation is asymptotically homophilic if $W_{ij} = \mathds{1}_{\eta_{ij} \leq w_n(\Vert X_i - X_j\Vert)}$, $w_n(x) \geq g(s_n x)$, where $g: [0; \infty[ \rightarrow [0;1]$ is decreasing and $\lim_{n \rightarrow \infty} s_n = \infty$. \\ b) Network formation is regularly asymptotically homophilic if $W_{ij} = 1$ whenever $w_n(\Vert X_i - X_j\Vert) \geq \eta_{ij}$, $w_n(x) =g(s_n x)$, where $g: [0; \infty[ \rightarrow [0;1]$ is a decreasing function such that $0 < \int_{\mathds{R}^d} g(\Vert y\Vert) \ dy < \infty$, and $\lim_{n \rightarrow \infty} n {s_n}^{-d} = \infty$.

Part a) formalizes the idea of asymptotic homophily; part b) provides regularity conditions for asymptotic results. The definition is new and the resulting formation mechanism is consistent with the empirical regularities of social networks jackson2010social: sparsity\footnote{It is natural to let the average degree of a node be constant or grow only slowly for most applications. For instance, the average number of friends is typically viewed as constant or slowly increasing as the network expands, requiring the probability of forming a link to decrease with the size of the network. Letting the degree increase, albeit slowly, allows one to take advantage of asymptotic approximations.}, transitivity/clustering\footnote{Intuitively, clustering occurs because groups of similar individuals tend to form connections. Moreover, as shown in the appendix, the clustering coefficient does not vanish asymptotically, in contrast to Poisson random graphs erdHos1960evolution or configuration models bender1990asymptotic.}, degree heterogeneity\footnote{Because of the influence of covariates, the expected degree varies across individuals. The underlying density affects the degree distribution because people with common characteristics have an easier time forming connections. Another source of degree heterogeneity can be accommodated for by adding expansiveness factors as previously discussed -- for instance by refining the sequence $s_n$ with pair-specific rates using the multiplicative heterogeneity model --, although this is not pursued here for simplicity.}, and homophily.

Asymptotic homophily formalizes the notion that homophilic behavior is pronounced relative to sample size in the spirit of a drifting sequence\footnote{Similarly to bekker1994alternative, who analyzes the behavior of IV estimators with many instruments, or borusyak2022quasi, who analyze shift-share instruments with a growing number of shocks. “The sequence is designed to make the asymptotic distribution fit the finite sample distribution better. It is completely irrelevant whether or not further sampling will lead to samples conforming to this sequence” bekker1994alternative.}; the degree to which individuals are selective is, in a sense, preserved as we proceed to an asymptotic approximation.

Comparison groups

Under an asymptotically homophilic network formation process, it is possible to derive estimators whose bias is asymptotically negligible, even if some confounders are unobserved. Given a comparison group $\mathcal{C}_i$ for individual $i$, I define a CATE\footnote{Note that $x_i$ is not fully observed. The CATE of a given individual is identified, but not the underlying function of $x$. For this reason, the CATE is mainly interesting on its own when one cares about the treatment effect of a specific unit.} estimator

equation[equation omitted — 291 chars of source]

where $\mathcal{C}_{it} \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \mathcal{C}_{i} \cap \{j\vert T_j = t\}$. Note that the comparison group will vary with the sample size, although the dependence is left implicit in the notation.

remarkIn what follows, the vector $X$ from the unconfoundedness condition is assumed to be part of the network formation for ease of exposition. In practice, the binding restriction is that $X^u$ belongs to it: If some observed covariate ($x_k$) that affects the outcome does not influence network formation, it can be controlled for nonparametrically by multiplying each weight $\mathds{1}(j \in \mathcal{C}_{it})$ with kernel weights $K_b(\Vert X_{ik}^o - X_{jk}^o\Vert)$ throughout. Because the weights based on the network act as (noisy versions of) kernel weights, the use of such hybrid weights is naturally nested in the current framework.

The first idea is to rely on friends to construct a comparison group, , i.e., set $\mathcal{C}_i=\mathcal{N}(i)$. This generates a CATE estimator based on the difference between treated friends and non-treated friends. Thanks to triangular inequality relationships and the nature of homophily, however, one can extract additional information from the friendship network. For instance, one can consider friends of friends or higher-order friendships:

align[align omitted — 347 chars of source]

for some upper order of friendship $M$. For $M=1$, this is the simple estimator that compares treated friends and non-treated friends. Although the estimator averages more observations as $M$ increases, the comparison group increasingly selects observations whose characteristics differ from those of $i$. $M=1$ can be a reasonable choice if the ATE is the target, but higher values can be useful, e.g., to estimate a specific CATE.

An alternative estimator relies on having at least $c$ friends in common:

align[align omitted — 437 chars of source]

Although both estimators allow for consistent estimation of treatment effects, the composition of the underlying comparison groups can differ significantly. Therefore, it can be a useful robustness check to compare their treatment effect estimates, for instance if one is concerned about peer effects or related issues that would likely affect these estimators in different ways.

As an illustration, consider Figure 1 where the comparison group for individual $i$ with unobserved characteristics $x_i \in \mathds{R}^2$ is shown in red and the remaining observations in black. The leftmost picture is the unknown\footnote{Except in the extreme network formation process in which individuals select friends deterministically conditional on covariates: $w(x) = \mathds{1}_{[0, C]}(x)$. Then, the first two pictures become identical.} group that consists of all observations below a certain distance. Next, on the right, friends are used as the comparison group, providing a noisy version of (i): selected observations tend to fall close to $x_i$, but some close observations are ignored while observations farther away may be selected nonetheless.

In the third picture, one looks at friends of friends. This allows us to make use of more observations, which will reduce the variance of the estimator, but there is also a tendency to grab more observations outside the sphere. Finally, the last picture selects individuals who have at least two friends in common with $i$. This typically reduces the bias compared to using the previous groups, but selects fewer observations.

figure[figure omitted — 526 chars of source]

CATE inference

Under asymptotic homophily, standard asymptotics are achievable with estimators such as ((ref)) or ((ref)), whose bias asymptotically disappears. As for the variance, its collapse is based on averaging over increasingly many observations. The asymptotic behavior of the size of a comparison group depends on the speed of convergence of $w_n$, of which the following lemma provides a formal analysis.

lemmaThe number of connections for $i$ exceeds any real number with probability approaching one if $\lim_{n \rightarrow \infty} n \int_{\mathds{R}^d} w_n(\Vert x_j - x_i\Vert) f(x_j) \ dx_j = \infty$. \\ Under regular asymptotic homophily, this holds, and moreover \\ a) The probability of forming a connection of order up to $M$ is $O\left({s_n}^{-d} \left(n {s_n}^{-d}\right)^{M-1}\right)$. The number of connections for $i$ of order up to $M$ then exceeds any real number with probability approaching one. \\ b) If $n {s_n}^{-2d} \rightarrow 0$, the probability of $i$ and $j$ having at least $c$ friends in common is $O\left({s_n}^{-2d} (n {s_n}^{-2d})^{c-1}\right)$.

The lemma, proven in the appendix, provides conditions under which the sizes of potential comparison groups grow to infinity. It also specifies the rates at which the probabilities of forming a connection up to the $M$-th or having more than $c$ friends in common decrease under asymptotic homophily, and when all corresponding counts go to infinity. With overlap, the lemma ensures the number of treated and untreated connections both grow to infinity.

The condition in the lemma states that $w_n$ must not vanish too fast to ensure that connections are still being formed. The formal criterion analyses the integral $\int_{\mathds{R}^d}w_n(\Vert x\Vert) f(x+x_i) \ dx$, suggesting that link functions that do not depend on network size or vanish uniformly slowly enough such as $w_n(x) = {s_n}^{-1} g(x)$ induce unbounded friend counts.

Basically, individuals must not become too selective too quickly to ensure that they keep forming connections. This is natural in many networks (e.g., friendship network) as the expected degree is often viewed as only slowly increasing. In the asymptotic homophily framework, this means $n {s_n}^{-d}$ increases at a low rate.

Asymptotic homophily is not only consistent with a growing count, but also implies an improvement in matching quality that is absent when $w_n$ is constant or grows uniformly. Specifically, a sequence such as $w_n(x) = {s_n}^{-1} g(x)$ would stabilize the “posterior” distribution $f_{X_j \vert j \in \mathcal{N}(i)}$; it does not imply that people improve their average match in larger networks.

As a result, regular asymptotic homophily will be key in securing consistency properties. An important part in establishing these is the analysis of the bias

align*[align* omitted — 339 chars of source]

which will be shown to disappear under various conditions. Specifically, the following conditions on $w_n$ will ensure that the bias vanishes.

assumption[Hölder continuity of CATE and convergence of link function] a) $\operatorname{CATE}(x)$ is Hölder continuous with exponent $\alpha$ on a neighborhood of $x_i$, i.e. for any $x, y$ in the neighborhood $\Vert\operatorname{CATE}(y) - \operatorname{CATE}(x)\Vert \leq C \Vert y - x \Vert^\alpha$ for some $\alpha > 0$. \\ b) For some $\varepsilon_n \downarrow 0$, either $\mathcal{C}_i=\cup_{m=1}^M \mathcal{N}^m(i)$ and $\sum_{m=1}^M (w_n(\frac{\varepsilon_n}{m}))^m = o({s_n}^{-d})$, or $\mathcal{C}_i=\{j\vert \ \vert \mathcal{N}(i;j)\vert \geq c\}$ and $(w_n(\frac{\varepsilon_n}{2}))^c = o({s_n}^{-d})$.

Hölder continuity is a standard assumption that imposes a mild degree of smoothness in the CATE. Part b) of the assumption restricts the way $w_n$ converges to $1$; it requires a sufficiently fast convergence away from the origin.

Although consistency can be achieved under very weak conditions, the rates of convergence may be low and the conditions on $w_n$ may be hard to interpret. Assuming that the underlying functions of $X$ are sufficiently smooth implies a clean rate of $O({s_n}^{-2})$ for the bias under regular asymptotic homophily. In practice, I require existence of first-order derivatives:

assumption[Existence of Derivatives] The conditional expectation \\ $\mathds{E}[Y_i(t)\vert X_i=x]$ and the propensity score $p(x) \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \mathds{P}[T_i=1\vert X_i=x]$ are continuously differentiable.

Now, the consistency theorem reads

theorem[Consistency] a) Suppose $\lim_{n \rightarrow \infty} n \int_{\mathds{R}^d} w_n(\Vert x\Vert) f(x+x_i) \ dx = \infty$, and $\mathds{E}[Y_j(t)^2] < \infty$ for $t = 0, 1$. \\ Then, $\widehat{\operatorname{CATE}}(x_i; \mathcal{C}_i)$ is consistent for $\operatorname{CATE}(x_i)$ under Assumption 2.2. Moreover, the bias satisfies $\mathds{B}_i = O(\varepsilon_n^\alpha + {s_n}^d R)$ with $R=\sum_{m=1}^M w_n(\frac{\varepsilon_n}{m})^m$ and $R=w_n(\varepsilon_n/2)^c$, respectively, for $\varepsilon_n \downarrow 0$ as in Assumption 2.2. \\ b) If the network formation is regularly asymptotically homophilic, Existence of derivatives holds, and the covariate density has bounded second-order derivatives, then the estimators are consistent with bias $\mathds{B}_i = O({s_n}^{-2})$.

The theorem is proven in the appendix. The main difficulty in part a) is to derive an expression for the bias that can subsequently be bounded via homophilic assumptions and triangular inequalities. In the second part, the existence of derivatives allows the use of Taylor expansions and the derivation shares similarities with nonparametric kernel analysis, though the noisy matching through $w_n$ makes the problem non-standard.

To sum up, consistency is secured provided homophilic behavior is preserved as sample size grows and $w_n$ does not drop too fast with sample size. If $w_n$ does not decrease to zero (so that there is still an asymptotic bias due to different covariates) or does so too quickly (so that the friend count shrinks), the estimator is no longer consistent. Because mean degrees are typically not increasing quickly with network size, these situation are seldom empirically relevant. When the network is dense -- the probability of forming a link does not drop to 0 -- inference results can be obtained using the method of Subsection (ref).

I finalize the analysis with an asymptotic normality result: CATE estimators are asymptotically normal at $x_i$. Formally,

theorem[Asymptotic Normality] Suppose that Consistency assumptions hold and the potential outcomes have $2+\delta$ moments for some $\delta > 0$ (conditional on $X=x_i$). Then, for a bias $\mathds{B}_i$ at location $x_i$, the (conditional) asymptotic distribution reads \begin{equation*} \sqrt{\vert \mathcal{C}_i\vert} (\widehat{\operatorname{CATE}}(x_i; \mathcal{C}_i)-\operatorname{CATE}(x_i)-\mathds{B}_i) \overset{d}{\rightarrow} \mathcal{N}\left(0; V\right) \end{equation*} where $V = \frac{\mathds{V}[Y_j(1)|X_j=x_i]}{\mathds{P}[T_j=1\vert X_j = x_i]}+\frac{\mathds{V}[Y_j(0)|X_j=x_i]}{\mathds{P}[T_j=0\vert X_j = x_i]}$. \\ Moreover, $\mathds{B}_i$ is asymptotically negligible if $\operatorname{CATE}(x)$ is Hölder continuous with exponent $\alpha$ on a neighborhood of $x_i$ and one of the following holds: \\ (i) $\mathcal{C}_i=\cup_{m=1}^M \mathcal{N}^m(i)$ with $\sum_{m=1}^M w_n(\frac{n^{-\gamma}n}{m})^m = o({\lambda_n}^2 \frac{M}{n})$ for $\gamma > \frac{1}{2 \alpha}$ or \\ (ii) $\mathcal{C}_i=\{j\vert \ \vert \mathcal{N}(i;j)\vert \geq c\}$ and ${s_n}^d w_n(n^{-\gamma})^c = o({\lambda_n}^{-1/2})$ and $\gamma > \frac{1}{2 \alpha}$ or \\ (iii) Network formation is regularly asymptotically homophilic, Existence of Derivatives holds, the covariate density has bounded second-order derivatives, and $\frac{\sqrt{\lambda_n}}{{s_n}^2} \rightarrow 0$ for $\lambda_n \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} n {s_n}^{-d}$.

When the CATE itself is of particular interest, the theorem provides a way to perform standard inference. The variance can be estimated with the same kind of truncation methods and one would generally increase $M$, as long as the bias remains negligible, in order to decrease the variance. Given that the size of the comparison group increases with powers of $M$, including neighbors of up to the fourth order, but not much higher, usually makes sense.

The size of the comparison group grows at rate $\lambda_n$. If $\lambda_n$ grows too fast compared to $s_n$, the estimator is still consistent though inference requires handling the bias. In social networks such as friendship networks, it is often reasonable to let the average degree grow only slowly newman2018networks so that a slow growth of $\lambda_n$ such as $\ln(n)$ can be justified.

Finally, consistent estimation of treatment effects at given $x_i$'s suggests that inference about average effects is possible. This is the next result: the average treatment effect can be estimated under unobservable-robust unconfoundedness.

ATE inference

Given the last two theorems, one obtains an estimator $\widehat{\operatorname{ATE}}$ by averaging over a CATE estimator at all $x_i$. The resulting ATE estimator is then consistent and asymptotically normal under the theorems' conditions and regularity conditions.

Specifically, consider $\frac{1}{n} \sum_{i=1}^n \widehat{\operatorname{CATE}}(x_i, \mathcal{C}_i)$ and collect the terms involving $Y_i$ for each $i$ to obtain $\frac{1}{n} \sum_{i=1}^n \left(T_i \sum_{j \in \mathcal{C}_{i}} \frac{1}{\vert \mathcal{C}_{j1}\vert} - (1-T_i) \sum_{j \in \mathcal{C}_{i}} \frac{1}{\vert \mathcal{C}_{j0}\vert}\right) Y_i$.\footnote{This last expression holds as long as $i \in \mathcal{C}_i$ implies $j \in \mathcal{C}_i$, which is true for the comparison groups previously discussed.} More compactly, the estimator reads

equation[equation omitted — 175 chars of source]

where $\theta_{i} = \sum_{j \in \mathcal{C}_{i}} \frac{1}{\lvert \mathcal{C}_{j T_i}\rvert}$.\footnote{Although the weights are well defined with probability approaching one, it may be useful to regularize them in finite sample by adding a small offset to the denominators.} Its asymptotic distribution is described by the following theorem:

theoremSuppose that $\mathds{E}[Y_j(t)^2] < \infty$, that condition (iii) of Theorem 2.2 holds, and that $\frac{\sqrt{n}}{{s_n}^2} \rightarrow 0$. Then, \begin{equation*} \sqrt{n} (\widehat{\operatorname{ATE}} - \operatorname{ATE}) \overset{d}{\rightarrow} N\left(0; \mathds{E}\left[\frac{\mathds{V}[Y_i(1)\vert X_i]}{p(X_i)} + \frac{\mathds{V}[Y_i(0)\vert X_i]}{1-p(X_i)}\right] + \mathds{V}[\operatorname{CATE}(X_i)]\right) \end{equation*}

The theorem is proven in the appendix. It can be seen that the asymptotic variance reaches the semiparametric efficiency bound for the estimation of the ATE (hahn1998role, hirano2003efficient). The theorem enables inference about the average treatment effect under (unobservable-robust) unconfoundedness in an asymptotically efficient way.

Testing for confounding effect of unobservables

Assuming unconfoundedness conditional on observed covariates $X_i^o$, suppose that a researcher uses observables to construct an estimator of the ATE, say $\widehat{\mbox{ATE}}^o$, whose influence function is the efficient score.

It may be of interest to test for a possible confounding effect of unobservables featured in network formation by testing whether the difference between $\widehat{\mbox{ATE}}$ and $\widehat{\mbox{ATE}}^o$ is statistically significant.\footnote{I thank an anonymous referee for suggesting this setup for a Hausman test.} Formally, if the regularity for asymptotic normality are met, then under the null that the outcome is unaffected by the unobserved components, i.e., $\mbox{CATE}(X_i) \equiv \mbox{CATE}(X_i^o)$, we can establish

align*[align* omitted — 441 chars of source]

which forms the basis for a test.

Dense networks with homophily in the unobservables

This section considers a link formation model that only assumes homophilic behavior in unobservables and is suitable for dense networks. The method provides insight about how to create sub-groups that are increasingly close in terms of unobservables and can be adapted to deal with empirical concerns such as matching on treatment status. Although dense networks are less frequent, this approach is thus valuable as it covers additional network structures and applications of interest.

Now, a link exists between $i$ and $j$ if $\eta_{ij} \leq w(h(X_i^o; X_j^o) + \Vert X_i^u - X_j^u\Vert)$. The function $w$ does not explicitly depend on $n$ anymore, although some dependence could be accounted for. The function $h$ need not be homophilic nor separable in the observed $X^o$ but is assumed known. Most results would apply with minor modifications if a lower bound with the relevant properties can be obtained. The dimension of $\mathcal{X}^u$ is denoted by $d_u$.

When observables affect the outcome, they can be controlled for using standard methods.\footnote{For instance, by constructing again products with kernel weights or performing regression adjustments.} For ease of exposition, I focus on presenting how to control for the unobserved components.

I consider again $\frac{1}{\vert \mathcal{C}_{i1}(\kappa)\vert} \sum_{j \in \mathcal{C}_{i1}(\kappa)} Y_j - \frac{1}{\vert \mathcal{C}_{i0}(\kappa)\vert} \sum_{j \in \mathcal{C}_{i0}(\kappa)} Y_j$, where the comparison group now depends on a truncation parameter $\kappa$. It will play a key role by placing a lower bound on $h$, inducing a closer distribution of unobservables among friends.

The main idea is that if there is no observed rationale for two people being friends, it becomes more likely that there is a unobserved reason for their friendship. Then, people that are friends despite a high value of $h$ are less likely to differ strongly on unobservables. To see this, consider for illustration the case where people reject friendship with anyone whose quality of match falls below a certain threshold (i.e., $w(x)=1$ for $x$ large enough). Then, two friends with an $h$ close to the boundary must have close unobservables since a high discrepancy in unobservables would have brought $h+\Vert X_i^u-X_j^u\Vert$ above the threshold.

This suggests using comparison groups of the form $\mathcal{C}_i(\kappa) \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \{j \in \mathcal{N}(i)\vert h(x_i^o, x_j^o) > \kappa\}$. The estimator then truncates the sums to select individuals whose observed characteristics make them unlikely to be friends. $\kappa$ is viewed as a sequence converging to $\infty$ at a rate to be determined. Using this strategy, a counterpart to Theorem 2.2 can be established with $\kappa$-truncation replacing asymptotic homophily.

theoremSuppose $\operatorname{CATE}(x)$ is Hölder continuous with exponent $\alpha$ on a neighborhood of $x_i$, $w$ has bounded support, and there exist a sequence $\lambda_n \rightarrow \infty$ and a sequence $b_n$ such that $\kappa$-truncation satisfies $n w(\kappa+b_n) b_n^{d_u+1} = O(\lambda_n)$ and a sequence $\varepsilon_n \downarrow 0$ satisfying $\kappa+\varepsilon_n > \sup\{\mbox{supp}\{w\}\}$ eventually. Then, the estimator satisfies \begin{equation*} \sqrt{\vert \mathcal{C}_i\vert} (\widehat{\operatorname{CATE}}(x_i; \mathcal{C}_i(\kappa))-\operatorname{CATE}(x_i)-\mathds{B}_i) \overset{d}{\rightarrow} \mathcal{N}\left(0; V\right) \end{equation*} and the bias is negligible if $\sqrt{\lambda_n} {\varepsilon_n}^\alpha \rightarrow 0$.

The condition on $w(\kappa+b_n) b_n^{d_u+1}$ restricts the speed at which $\kappa$ can increase so that the number of observations used in estimating the CATE keeps growing. The first part pertains to the behavior of the $w$ function; the second term pertains to the space in which unobservables live.

The term $w(\kappa+b_n)$ comes from the increasing cost of truncating as potential connections are accepted at decreasing rates. In the presence of a discontinuity at the end of the support, i.e., $w(x) = a \mathds{1}_{x \leq D}$ for $a \in ]0; 1]$ and $D\in \mathds{R}^+$, this term disappears. The second term, ${b_n}^{d_u+1}$, is the result of forcing unobservables in a $b_n$-ball using values of $h$ lying between $\kappa$ and $\kappa+b_n$.

A natural estimator of the end of support, if unknown, is the highest value of $h$ among all $i, j$ satisfying $W_{ij}=1$. Unbounded support of $w$ can be accommodated, provided the function vanishes sufficiently quickly (the threshold is exponential: functions decaying faster than $w=e^{-x}$ provide a sufficiently fast decay). In this case, a factor of $\rho_h(\kappa + b_n/2)$, where $\rho_h$ is (a lower bound on) the tail decay of $h$ conditional on $X_j^u \in B_{r}(x_i^u)$ for some $r$, has to be added. The reason is the tail decay of the density while one seeks increasingly larger values of $h$.

The conditions on $\varepsilon_n$ ensure that the bias disappears sufficiently fast to enable standard inference. A more primitive statement is $w(\kappa + \varepsilon_n) = o(\lambda_n/n)$, which mirrors conditions such as those in Assumption (ref), part b); this boils down to $\kappa + \varepsilon_n$ eventually crossing the end of the support of $w$ when it is finite.

Finally, one can again consider alternative comparison groups -- for instance, by including friends of friends or people with sufficiently many common friends -- or construct the truncated group differently -- for instance, considering triangle of friends can give more leeway to vary $h$.

Overall, these results show that suitably refining a comparison group such as friends using the observed covariates allows one to isolate increasingly good matches in terms of unobservables. Although the levels of the unobserved variables is not identified, groups with increasingly similar values can be recovered. The rates are slower than those arising from asymptotically homophilic networks, though it should be noted that they focus directly on the unobserved components while the observed variables can be adjusted in more traditional ways, e.g., regression adjustments. This is sufficient to recover consistent estimators of treatment effects under unobserved confounders with dense networks and to learn who is comparable to whom in terms of unobservables.

Simulations

I assess the performance of the estimators through simulations. I consider various outcome equations and measure the resulting root mean square error (RMSE). The RMSE of standard estimators making use of observed variables is provided for comparison.

The variables are generated as follows: a random vector $V$, whose components are uniform, triangular, and sum of three uniforms, is used to construct

equation*[equation* omitted — 204 chars of source]

so that there is a non-trivial correlation structure among the components of $X$. In the baseline, $\mbox{corr}(X_1, X_2)=0.28$, $\mbox{corr}(X_1, X_3)=-0.26$, and $\mbox{corr}(X_2, X_3)=-0.32$, though the results exhibit similar patterns for weaker or stronger correlations. The variances are normalized to one. The first two variables are observed, but the last one is not, i.e., $X^o = (X_1, X_2)$ and $X^u = X_3$.

The propensity score follows a logistic distribution with argument $X \beta$, where $\beta =

pmatrix[pmatrix omitted — 30 chars of source]

'$, and the treatment status is then drawn conditional on $X$. The parameter $\beta_3$ controls selection on unobservables and takes value in $\{0, 0.5, 1\}$. In the first case, the probability of being treated does not change with $X_3$; in the last case, the unobserved variable is on a similar footing as each of the observed ones. The performance of traditional methods that cannot account for the unobserved component is expected to deteriorate as $\beta_3$ increases.

The outcome equation is given by $y = 5 + \mbox{CATE}(x) T + g(x) + \varepsilon$. I explore three specifications:

itemize• Homogeneous treatments effects ($\mbox{CATE}(x) \equiv 1$) with linear impact of unobservables ($g$ is linear in $x$) • Heterogeneous treatments effects ($\mbox{CATE}(x) = 2 \Phi(-x_1+x_2+x_3)$) with linear impact of unobservables ($g$ is linear in $x$) • Heterogeneous treatments effects ($\mbox{CATE}(x) = 2 \Phi(-x_1+x_2+x_3)$) with quadratic impact of unobservables ($g$ contains both a linear term in $x$ and a quadratic term ${x_3}^2+x_2 x_3$)

The network formation process uses $w(x)=e^{-\frac{1}{2} x^2}$.\footnote{Other specifications such as $w(x)=\mathds{1}_{x<1}$ or $w(x)=\max\{1-x, 0\}$ deliver similar results.} The baseline sample size is $n=500$ and $s_n$ is calibrated so that the average number of friends is roughly five-six, which is around the number of close friends people often report, e.g., in the application. As $n$ rises, $s_n$ evolves at the rate $n / \ln(n))$.

The researcher controls for observables using kernel weights that multiply network weights, as previously described. People form friendship links based on $X_2$ and $X_3$. As a result, only $X_1$ is unaccounted for in network weights, but a researcher observing $X_2$ may want to further include it in the kernel weights. I consider both possibilities (referred to as base- and over-controlling case below). Alternative methods (e.g., OLS) always make use of all observables.

The results for all estimators and sample sizes $n=500, 2000$ are reported in Appendix C. For a simple and representative review, the results for the specifications $M=1, c=2$ (orange and yellow lines, respectively) and OLS (blue line) are graphically depicted below for the 3 types of outcome equations and $n=500$.

figure[figure omitted — 510 chars of source]

In the linear homogeneous case, the OLS estimator performs well, especially when there is no selection on unobservables. Its performance deteriorates as $\beta_3$ increases and it reaches an RMSE similar to that of the proposed estimators when the selection on the unobserved variable is on par with the level of selection on observables. Furthermore, it behaves poorly when the unobservables become more prevalent in the functional form or the treatment response.

In contrast, estimators that leverage network information are able to perform similarly regardless of the strength of selection in unobservables and dominate OLS across specifications unless we force homogeneity and linearity. In the presence of unobserved confounding, the properties of OLS become quickly unappealing, especially when the unobserved variable has nonlinear impacts on the outcome or the treatment effect. Other methods based on estimating the propensity score also fail to deliver reliable estimates of the ATE as they can only accommodate heterogeneity and nonlinearities that arise from observed confounders. As a result, they are even dominated by OLS and a fortiori by the proposed estimators.

Finally, the simulations provide some guidance for empirical decisions. The choice of $M$ or $c$ does not appear to have a strong influence on the performance of the estimator. Over-controlling leads to a general increase in RMSE, but the estimators still perform well and improve substantially over OLS in nonlinear or heterogeneous settings. Therefore, it is advisable to separately control for observables whose relevance to network formation is uncertain.

Application

I provide an application of the method to the estimation of the effect of parental involvement on students' test scores. I use the dataset from the project “Attitudes and Relationships among Primary and High School Students”, see DVN/ZHCTCK_2022. The dataset contains information about 4409 Brazilian high-school students, their beliefs, and friendship ties among them.

The outcome of interest is the average grade, ranging from 0 to 10, and the treatment is the level of parental support. The average is taken over math, Portuguese, English, history, geography, and art grades.\footnote{Although additional subjects are available, including them would require dropping a substantial number of observations because of missing data.} For comparability across school years, I normalize the grade by subtracting the mean over students in a given year. The dataset contains a (self-reported) score for parent support which ranges from around -3 to 1; the treatment is dichotomized by truncating around the mean of 0 and the goal is to estimate the ATE.

Although some covariates are available -- age, gender, race, religion, class dummies, whether parents are employed, poverty, and importance of study for the child, -- they are at best proxies for the underlying causes. Thus, omitted variable bias remains a concern, especially due to the usual unobserved ability.

The sign of the omitted variable bias is unclear in this context, even assuming that the score for the importance of studying controls adequately for motivation. Intuitively, children with lower ability may require more attention but are also more likely through heredity to have less able parents, who may be less inclined or able to help.

Formally, a possible model for the individual's grade, $y_i$, would be $y_i = F(a_i; m_i; s_i)$ where $F$ is an unknown function, $a_i$ is ability, $m_i$ is the level of (intrinsic) motivation\footnote{Grades may be affected directly or indirectly through increased effort. The target is the total effect that includes the mediated effect, conditioning on intrinsic motivation and ability.}, and $s_i$ is parental support. Denoting parents' ability by $A_i$, one could specify $a_i = A_i+\mbox{noise}_i$ and $s_i = \delta_1 a_i + \delta_2 A_i + \mbox{noise}_i$ with independent noises. This would imply $s_i = (\delta_1+\delta_2) a_i + \mbox{noise}_i$, where the signs of the deltas are likely to be negative and positive, respectively, leading to an indeterminate correlation sign. Hence, it is not obvious whether parents pay more attention to children with higher or lower ability. Moreover, the function $F$ is unlikely to be linear since, e.g., ability may increase the return to motivation and parent support and there may be nonlinear returns to motivation and/or ability.

As a result, not only is the OLS estimate likely to be unreliable, but it is also difficult to figure out the direction of the bias. This motivates the use of alternative methods that can account for the presence of unobservables.

There is evidence that people, teenagers in particular, form friendship links based on intelligence. Some studies clark1992friendship, burgess2011school, boutwell2017general have documented homophilic matching on various measures of intelligence among teenagers. According to boutwell2017general, “preadolescent friendship dyads are robustly correlated on measures of general intelligence”.

It is thus plausible that friendship ties account for ability, suggesting an avenue for correcting the ability bias through the estimator developed in this paper.\footnote{There is also some evidence for homophily in some of the covariates. The ratio of the average distance in age and gender among friends to the average distance between any two individuals is about one fourth and one half, respectively.} For each individual, I use friends as a comparison group\footnote{The number of people with no reported friend is relatively high, but most of these students are also missing covariates. This suggests that zero friend counts are more indicative of a missing data problem than of general asocial behavior. Consequently, I restrict the sample to the sample used for OLS, which also facilitates comparability. The effective sample size is then 2777 and people have an average number of about 5 friends, as calibrated for the simulation study. This relatively low number is consistent with situations in which people report close friends, which is generally the preferred target for the type of matching exercises that the method exploits.} and compute $\widehat{\operatorname{ATE}}$.

One may also consider controlling for all observed variables for further comparison to OLS or because it is believed that some variable, e.g., the importance of study score, should be directly controlled for. Since they are quite numerous, a possibility is to perform regression adjustments as follows: run a regression of $y_i$ on the treatment and relevant controls for each $i$ using observations in $\mathcal{C}_i$, then form an optimal weighted average of the treatment effect estimates.

According to OLS, the treatment has a small effect of 0.04. The estimator developed in the paper, however, suggests a much higher effect of $0.21$. Using all controls included in the OLS regression yields a treatment effect estimate of $0.14$ (standard error $0.07$). Because of the control for unobserved ability and its potential nonlinear interactions, the latter estimate may be more reasonable.

table[table omitted — 371 chars of source]
remarkRunning a regression without the parent's employment status and the poverty score increases the treatment effect estimate from OLS to 0.11. Going back to the model $s_i = \delta_1 a_i + \delta_2 \mbox{parents' ability}_i + \mbox{noise}_i$, this may indicate that OLS may be more biased upon controlling for parent's employment and poverty because these variables may act as proxies for parent ability and thus increase the conditional correlation between ability and parent support.
remarkAnother sanity check consists in varying the composition of the comparison group. Adding friends up to the second order or using people with at least two friends in common as a comparison group produces similar results: the effect is estimated to be $0.21$ or $0.18$. Because these versions of the estimator rely on comparison groups with different relationships to the individual, this can alleviate concerns that the effect is driven by other factors.