The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
72,675 characters
Party On: The Labor Market Returns to Social Networks in Adolescence
\title{Party On: The Labor Market Returns to Social Networks in Adolescence\thanks{We are grateful to seminar and conference participants at Northwestern
University, Santa Barbara University, USC, Peking University, SOCCAM,
the All-California Labor Economics conference, the North American
Winter Meeting of the Econometric Society, the Asian Meeting of the
Econometric Society in China, and the China Labor Economist Conference.
We are particularly grateful to Larry Katz for his very insightful
suggestions. This project was supported by the California Center for
Population Research at UCLA (CCPR), which receives core support (P2C-
HD041022) from the Eunice Kennedy Shriver National Institute of Child
Health and Human Development (NICHD). Research done prior to joining
Audible. Audible is a subsidiary of Amazon. All views contained in
this research are the author's own and do not represent the views
of Amazon or its affiliates. All errors are our own.}}
\author{Adriana Lleras-Muney, Matthew Miller, Shuyang Sheng, and Veronica
Sovero\thanks{Adriana Lleras-Muney, Department of Economics, UCLA, email: [email removed];
Matthew Miller, Audible Inc, email: [email removed]; Shuyang Sheng,
Department of Economics, UCLA, email: [email removed]; Veronica
Sovero, Department of Economics, UC Riverside, email: [email removed].}}
\maketitle
\begin{abstract}
\begin{singlespace}
\noindent We investigate the returns to adolescent friendships on
earnings in adulthood using data from the National Longitudinal Study
of Adolescent to Adult Health. Because both education and friendships
are jointly determined in adolescence, OLS estimates of their returns
are likely biased. We implement a novel procedure to obtain bounds
on the causal returns to friendships: we assume that the returns to
schooling range from 5 to 15\% (based on prior literature), and instrument
for friendships using similarity in age among peers. Having one more
friend in adolescence increases earnings between 7 and 14\%, substantially
more than OLS estimates would suggest.
\end{singlespace}
\end{abstract}
\begin{singlespace}
\section{Introduction}
\end{singlespace}
An individual's social capital (the number and quality of their connections)
impacts many economic outcomes. Individuals' classroom and school
peers impact their education (\citealp{Sacerdote2001}; \citealp{Carrell2009}),
earnings (\citealp{carrell2018disruptivepeers}; \citealp{Michelman2021})
and health (\citealp{Carrell2011}). Whom one befriends from among
these peers also influences important lifetime outcomes: \citet{Chetty2022}
show that the number of high SES friendships and economic mobility
in a neighborhood are positively associated. However, there is little
evidence on whether the number of one's friends (not just the types
they are exposed to) causally affect labor market outcomes.
Using data from the National Longitudinal Study of Adolescent to Adult
Health (hereafter Add Health), which follows individuals from adolescence
into adulthood, we study the labor market returns to adolescent friendships.
During adolescence, individuals develop important cognitive and social
skills, have key educational outcomes determined and form friendships
\citep{heckmanmobility2014}. Friendships are associated with better
larbor market outcomes: Figure \ref{fig:returns_edu_friend} plots
log annual earnings at ages 24--34 against the number of friends
at ages 12--20, separately for males and females. The non-parametric
lines show a striking positive association between adolescent friendships
and young adult earnings for both sexes.\footnote{We detail later in the paper how we measure the number of friends.}
OLS estimates of this association may be biased downward however:
individuals with social skills may engage in activities excluded from
regressions, like drinking and partying, that help them make friends
but reduce their education and earnings. This ``taste for partying''
is unobserved and omitted from earnings regressions, and can cause
the return to friendships to be downward biased. Conversely, social
individuals may prefer productive social activities (working together),
which improve their networks and yield higher earnings, biasing the
return to friendships upward.\footnote{Social skills are associated with large and growing returns in the
labor market (\citealp{Weidmann2021}; \citealp{Deming2017}).}
To estimate the causal returns to friendships, we construct an instrument
for the number of friends that exploits the fact that homophily (similarity
in traits) predicts friendship formation (\citealp{Boucher2015};
\citealp{Jackson2014}). Our instrument is the average absolute difference
in age between a student and the peers in her school and grade, which
we refer to as age distance. Because our baseline regression controls
for own age and the average age of individuals in the school-grade,
as well as grade and school fixed effects, age distance will vary
across individuals because the distribution of ages varies across
schools and grades. For example, a 13.5 year old with two peers aged
12.5 and 14.5 has an age distance of 1 while mean age is 13.5. If
the same 13.5 year old were instead in a group with a 13 year old
and a 14 year old, mean age would be the same, but age distance would
be smaller (0.5), resulting in more friends. This variation is similar
to what is used in \citet{Bifulco2011classmates} and \citet{carrell2018disruptivepeers},
who leverage variation across cohorts within schools to estimate the
effects of peer characteristics on outcomes.\footnote{These papers typically investigate the direct (reduced form) relationship
between peer traits and outcomes. Although we leverage similar (though
not identical) variation, we use this variation in peer characteristics
as an instrument for friendships -- we do not study the direct effect
of peer similarity on outcomes.} In our setting, there is also variation in the instrument \textit{within}
school-grade because age distance depends both on the mean age in
their group and the student's own age. Our identification strategy
is most similar to \citet{Fletcher2020} who investigate the effect
of friendships on education outcomes using similarity in race/ethnicity
and socioeconomic status as instruments for friendships. We investigate
the effect of friendships on earnings instead.
Education is also endogenous in these regressions because educational
and social investments are determined jointly in adolescence. Studies
with well-identified causal returns to education make use of large
data sets and typically exploit state-level variation in compulsory
schooling or the cost of schooling, such as distance to school (\citealp{Mountjoy2022})
or school openings (see \citet{card2001estimating} for a review and
\citet{Oreopolous2020} and \citet{Psacharopoulos2018} for more recent
examples). In our data, these instruments are weak predictors of education.
We propose a novel approach to estimating the returns to friendships
that does not require instruments for education. We rely on the findings
from the previous literature and assume that the causal returns to
a year of schooling range from 5 to 15\% (\citealp{Oreopolous2020};
\citealp{Psacharopoulos2018}). Under this assumption, and making
use of the homophily-based instrument for friendships, we derive bounds
for the returns to friendships.\footnote{We are grateful to Larry Katz for suggesting this approach.}
We find that the returns to having one more friend during adolescence
range between 7 and 14\%, similar to the returns to one more year
of schooling. Our identifying assumption is that age distance determines
earnings only through its effects on education and friendships, conditional
on own characteristics, mean peer characteristics, and grade and school
fixed effects. The results are robust to a number of checks, including
addressing the potential concern that parents sort into schools or
grades based on relative age. These instrumented returns to friendships
are larger than OLS estimates. Computations suggest that measurement
error can explain the discrepancy, consistent with \citet{Griffith2021}.
The result is also consistent with downward omitted variable bias:
the data show that preferences for activities like drinking are associated
with more friendships but lower GPA and earnings.
We contribute to the well-established literature examining peer effects
among adolescents and young adults (see the review by \citealp{Jeon2015}).\footnote{Identifying how the behaviors and characteristics of peers affect
individuals was first discussed by \citet{Manski1993}. A few early
papers (\citealp{BRAMOULLE2009identification}) attempted to use friendship
networks to instrument for peer outcomes and overcome the joint determination
of outcomes (the reflection problem). These papers assume that a friendship
network is exogenous or endogenous through unobserved group heterogeneity.} For example, having peers with good academic outcomes improves one's
academic outcomes (\citealp{Carrell2008,Carrell2009}; \citealp{SACERDOTE2011peer};
\citealp{Bifulco2011classmates}; \citealp{Denning2021}). Peer effects
in adolescence can also carry into adulthood. \citet{carrell2018disruptivepeers}
find that having disruptive peers in the classroom during adolescence
has deleterious effects on individuals' labor market outcomes as adults,
partly because disruptive peers lower test scores.
Our paper suggests a new mechanism by which peer characteristics might
operate: friendship formation. A few papers investigate the effects
of friendship networks on educational outcomes using data on friendship
nominations (\citealp{Babcock2008}; \citealp{LavySand2019}; \citealp{Fletcher2020}).
However, friendships may also matter for labor market outcomes, separately
from their effects on education attainment. Previous work documents
that the size and connectedness of an individual's social network
in adulthood can improve wages and job match quality (\citealp{granovetter1973strength};
\citealp{montgomery1991social}; \citealp{calvo2004effects}; \citealp{Cappellari2015};
\citealp{Dustmann2015}). In addition to the number of friends, the
types of friends one is associated with may also impact labor market
outcomes (\citealp{Chetty2022}). Although our results on this are
only suggestive (we do not have powerful instruments for different
types of friendships), we find that almost all types of friendships
have returns in the labor market, including friends from low SES backgrounds
and with weak connections. This suggests that the overall number of
friendships -- not just their type -- is a relevant factor in determining
earnings.
There are no papers we are aware of that estimate the \textit{causal}
returns to the number of friendships on labor market outcomes --
this is the main contribution of this paper. The closest paper to
ours is by \citet{conti2013popularity}, who estimate returns to high
school friendships on wages using data from the Wisconsin Longitudinal
Survey. Their estimation strategy corrects for non-classical measurement
error in the number of friendships due to under-sampling, but they
do not account for endogeneity in both education and friendships which
this paper addresses.
Our second contribution is methodological. A strand of econometric
literature considers models with \textit{endogenous} networks (\citealp{goldsmith2013social};
\citealp{Hsieh2015network}; \citealp{Badev2021}; \citealp{johnsson2021peer};
\citealp{Auerbach2022}; \citealp{Griffith2022}; \citealp{Sheng2023}).
But these papers do not consider the joint determination of education
and friendships. As a result, they do not address how the endogeneity
in education complicates the identification of causal labor market
returns to friendships.
Our identification approach is motivated by the econometric literature
on partial identification, which is widely used as a remedy for identification
failure in various applications such as interval data \citep{Manski2002},
missing data \citep{Manski2003}, auctions \citep{Haile2003}, and
games with multiple equilibria \citep{Tamer2003}. We exploit the
idea of partial identification to resolve a new challenge where we
lack a good instrument for one of the endogenous regressors, but the
coefficient of that regressor is not of primary interest and can be
bounded. To our knowledge, this is the first paper that proposes a
partial identification approach to avoiding instrumenting for a secondary
endogenous regressor.
\section{Conceptual Framework: How are Friendships Formed in Adolescence?}
We start by summarizing a basic model of education and friendship
formation (Appendix \ref{app:Model}) and describe its implications
for the empirical analysis.\footnote{All figures and tables designated with a letter (e.g., \textquotedblleft A\textquotedblright ,
``B'') are shown in the Online Appendix.}
During adolescence, individuals decide how to allocate their time
between studying, socializing, and leisure. Studying increases educational
attainment, and socializing increases one's number of friends, both
of which have positive returns in the labor market. In deciding how
to allocate their time, individuals consider the returns to each activity,
which depend on their innate intelligence and social skills (Proposition
\ref{prop:endow}).
Because time spent investing in education and friends is determined
at the same time, both education and friendships are potentially endogenous
in a Mincer earnings equation and thus OLS estimates of their returns
may be biased. However, the sign of the bias in friendship returns
is not clear ex-ante (Proposition \ref{prop:ols}). On the one hand,
social skills may determine the number of friendships and have an
independent effect on earnings, causing an upwards bias in the returns
to friendships. On the other hand, partying and drinking may increase
friendships but negatively affect skill accumulation and labor market
outcomes, causing a downward bias in the returns to friendships.
To account for the endogeneity of friendships, we will exploit the
fact that an individual's accumulated social capital also depends
on the traits and decisions of their school-grade peers because the
production of friendships requires coordination with others -- you
cannot ``party alone''. Conditional on time spent socializing, individuals
that are more similar to each other are more likely to become friends.\footnote{In equilibrium this may not hold (see Proposition \ref{prop:homophily}).}
\section{Data}
We use the restricted-use National Longitudinal Study of Adolescent
to Adult Health (Add Health). The in-school sample is a complete census
of all students enrolled in a given school during the 1994--95 school
year. The in-school sample data include basic demographics as well
as friendship nominations. A random sample of the students interviewed
in school was selected for in-home interviews in Wave 1 during the
1994--95 school year (ages 12--20 years) and tracked over in subsequent
survey waves. In Wave 4 (which was conducted in 2008--09), respondents
were age 24 to 34, on average 29 years old. For this in-home sample,
we observe measures of endowments, investments, and cognitive and
social outcomes.
Our estimation sample is constructed from the in-home sample, but
we use information from the in-school sample to construct friendship
and homophily measures. We include 10,605 individuals with complete
data for gender, age, and race. Summary statistics for the estimation
sample are presented in Table \ref{tab:sum_stats}.
\textbf{Outcomes.} The main outcome of interest is total earnings
from wages or salary in the last year. If a respondent replied \textquotedblleft do
not know\textquotedblright{} to the earnings question, they were prompted
with twelve categories of earnings. We use the midpoint of the selected
range for these respondents (approximately 2\% of the sample). We
drop individuals who reported zero earnings (\textasciitilde 6\%),
which means they were unemployed the entire year. Individuals in our
sample made roughly \$38,000 in the previous year.
\textbf{Intermediate outcomes: friendships and education.} Our primary
measure of the number of friendships is a person's \emph{grade in-degree}.
For a given student, grade in-degree is the number of people \textit{within
the same school and grade} who nominate them as one of their friends
in Wave 1.\footnote{Individuals could list up to five nominations of each gender. }
In-degree has been widely used in the social network literature as
an objective measure of an individual's number of friendships because
it does not rely on self-reporting \citep{conti2013popularity}. Importantly,
we observe all nominations sent to students in the in-home sample
because the network data is derived from the in-school sample. Unlike
Add Health's measure of in-degree, we exclude nominations from students
in different grades: the majority of nominations (77\%) occur within
the same grade (on average, school in-degree is 4.4, whereas grade
in-degree is only 3.4 (Table \ref{tab:sum_stats})).
Education is measured by years of schooling. On average, individuals
in our data obtain almost 15 years of schooling. We also observe GPA,
a measure of how much students learned in school.
\textbf{Endowments}. We use self-reported extroversion, collected
in Wave 2, as the main measure of the social endowment of individuals.
Extroversion is one of the ``Big Five'' psychological traits. About
65\% of individuals report being extroverted.\footnote{The survey question is ``You are shy?'', and the the choices are
``strongly agree / agree / neither agree nor disagree / disagree
/ strongly disagree''. Individuals choosing last three categories
are defined as extrovert. Due to the survey design this measure is
missing for 26\% of individuals in the data. To maximize sample size,
we impute this measure and include a dummy for whether it is missing.} Consistent with existing evidence (\citealp{lenton2014personality}),
extroverts have larger earnings, and perhaps not surprisingly, more
friends (Columns 2 and 3 of Table \ref{tab:sum_stats}). Most interestingly,
they also have more years of schooling.
The Add Health Picture Vocabulary Test (AHPVT) score is our main measure
for cognitive endowments. This test, administered in Wave 1, is an
abbreviated version of the widely used Peabody Picture Vocabulary
test and measures verbal ability. While it is not an overall measure
of intellectual ability, it has a high correlation with other intelligence
tests (\citealp{Hodapp1999}; \citealp{Dunn2007}). For simplicity,
we refer to it as IQ. As expected, individuals with above median IQs
have greater earnings and education (more years of education and higher
GPA). Perhaps surprisingly, they also have more friends (Columns 4
and 5 of Table \ref{tab:sum_stats}).
\section{Empirical Strategy}
We now turn our attention to estimating the labor market returns to
friendships. We follow the previous literature and estimate the following
earnings equation (conditional on employment):
\[
Y_{i}=r_{e}E_{i}+r_{f}F_{i}+\beta'X_{i}+\gamma'\bar{X}_{gs}+\alpha_{g}+\lambda_{s}+\epsilon_{i},
\]
where the outcome of interest is the log annual earnings $Y_{i}$
in Wave 4 for a given individual $i$ (observed in grade $g$ and
school $s$ during Wave 1), $E_{i}$ stands for years of schooling,
and $F_{i}$ is a measure of $i$'s number of friends (such as in-degree).
We control for $i's$ characteristics $X_{i}$ (age in Wave 1, sex,
race, IQ, and extroversion) and mean characteristics in $i's$ school-grade
$\bar{X}_{gs}$ (mean age, fraction female, fraction white, mean IQ
and fraction extrovert). We control for grade fixed effects ($\alpha_{g}$)
to account for differences across grades within a school. We include
school fixed effects ($\lambda_{s}$) to control for unobserved school-level
characteristics that could be correlated with labor market outcomes
and potentially sort individuals into schools. Standard errors are
clustered at the school level.
The object of interest is the coefficient $r_{f}$, measuring the
returns to having one more friend. Proposition \ref{prop:ols} shows
that OLS estimates of $r_{f}$ are biased because education and friendships
are jointly determined. We take an instrumental variable approach
to overcome the endogeneity issue, which will also address classical
measurement error in number of friends.
We use homophily measures as instruments for friendships, following
the evidence that individuals that resemble each other are more likely
to become friends \citep{Jackson2008}. Unfortunately, the instruments
for education used in the literature (quarter of birth, distance to
school) are weak in our data. Instead, we propose a novel approach
to estimating the returns to friendships that does not require an
instrument for education.\footnote{The education literature shows that relative age within a classroom
also affects educational attainment (e.g., \citealp{black2011too}).
However, we experimented using homophily measures for education as
well as for in-degree and found that homophily instruments fail the
weak IV tests if we use them to predict both education and friendships.}
\subsection{\label{sec:bounds}Bounding the returns to friendships}
There is a substantial literature estimating the (causal) returns
to schooling in the United States using various approaches. While
estimates differ across studies and populations, causal estimates
of $r_{e}$ typically lie between 5 and 15\% \citep{card2001estimating,Psacharopoulos2018,Oreopolous2020}.
Instead of estimating $r_{e}$, we assume that $r_{e}$, while unknown,
lies in this range. Then by instrumenting for friendships only, we
derive upper and lower bounds for the (causal) returns to friendships.
Let $\theta=(r_{f},\beta',\gamma',\alpha',\lambda')'$ denote the
vector of parameters, where $\alpha=(\alpha_{g},\forall g)'$ is the
vector of grade fixed effects, and $\lambda=(\lambda_{s},\forall s)'$
is the vector of school fixed effects. Let $\tilde{X}_{i}$ denote
the vector of friendships $F_{i}$ and other covariates (individual
characteristics $X_{i}$, school-grade mean characteristics $\bar{X}_{gs}$,
and grade and school dummies), and $Z_{i}$ the vector that consists
of the instruments for $F_{i}$ (homophily measures) and the covariates.
Ideal instruments for friendships satisfy two conditions: (i) they
predict friendships $F_{i}$ ($\mathbb{E}[Z_{i}\tilde{X}'_{i}]$ has
full column rank); and (ii) they are excluded from the earnings equation
($\mathbb{E}[Z_{i}\epsilon_{i}]=0$). Instruments for friendships
can also predict education. The exclusion restriction is satisfied
if the instruments do not affect earnings except through friendships
and education. This exclusion restriction implies the moment condition
\begin{equation}
\mathbb{E}[Z_{i}(Y_{i}-r_{e}E_{i}-\theta'\tilde{X}_{i})]=0.\label{eq:moment}
\end{equation}
Suppose that $W$ is a positive definite weighting matrix. If the
education return $r_{e}$ were known, the parameter $\theta$ would
satisfy $\theta=\mathbb{E}[Q_{i}(Y_{i}-r_{e}E_{i})]$, where $Q_{i}$
denotes the vector $(\mathbb{E}[\tilde{X}_{i}Z'_{i}]W\mathbb{E}[Z_{i}\tilde{X}'_{i}])^{-1}\mathbb{E}[\tilde{X}_{i}Z'_{i}]WZ_{i}$.
Because the true value of $r_{e}$ lies between $r_{e}^{l}=0.05$
and $r_{e}^{u}=0.15$, the parameter $\theta$ is bounded between
$\mathbb{E}[Q_{i}(Y_{i}-r_{e}^{l}E_{i})]$ and $\mathbb{E}[Q_{i}(Y_{i}-r_{e}^{u}E_{i})]$.
In practice, we can estimate the bounds by setting $r_{e}$ at $r_{e}^{l}$
and $r_{e}^{u}$ and estimating a GMM regression of $Y_{i}-r_{e}E_{i}$
on friendships $F_{i}$ and other covariates, using $Z_{i}$ as the
instrument. For example, if we set $r_{e}=r_{e}^{l}$ and regard $Y_{i}-r_{e}^{l}E_{i}$
as the dependent variable, then the GMM estimator yields an estimator
for the bound $\mathbb{E}[Q_{i}(Y_{i}-r_{e}^{l}E_{i})]$. The bound
$\mathbb{E}[Q_{i}(Y_{i}-r_{e}^{u}E_{i})]$ can be estimated similarly.\footnote{To ensure that the upper and lower bounds are estimated using the
same weighting matrix, we use the GMM estimator that assumes homogeneous
errors, that is, we set $W=\mathbb{E}[Z_{i}Z'_{i}]^{-1}$. This estimator
is equivalent to 2SLS.} For each component of $\theta$, we then obtain a consistent estimator
for the upper (lower) bound by taking the maximum (minimum) of the
two estimates of the component.\footnote{\label{fn:bound}Whether the upper or lower bound of a component of
$\theta$ is achieved at $r_{e}^{l}$ or $r_{e}^{u}$ depends on the
sign of the corresponding component of $\mathbb{E}[Q_{i}E_{i}]$.
For example, let $Q_{i,1}$ denote the first component of $Q_{i}$.
Then $r_{f}$ has the upper bound $\mathbb{E}[Q_{i,1}(Y_{i}-r_{e}^{l}E_{i})]$
if $\mathbb{E}[Q_{i,1}E_{i}]\geq0$ and $\mathbb{E}[Q_{i,1}(Y_{i}-r_{e}^{u}E_{i})]$
if $\mathbb{E}[Q_{i,1}E_{i}]<0$. The lower bound of $r_{f}$ can
be derived by swapping $r_{e}^{l}$ and $r_{e}^{u}$. In practice,
the components of $\mathbb{E}[Q_{i}E_{i}]$ can be recovered by regressing
education $E_{i}$ on friendships $F_{i}$ and other covariates, using
$Z_{i}$ as the instrument. The sign of the coefficient on friendships
determines whether the upper bound of $r_{f}$ is achieved when $r_{e}$
is high or low. A positive friendship coefficient implies that the
upper (lower) bound of $r_{f}$ is achieved at low (high) $r_{e}$.}
\textbf{Inference.} We follow \citet{Imbens2004} to construct a confidence
interval for any true value of $\theta$ that lies between the bounds.
These confidence intervals are asymptotically valid regardless of
whether $\theta$ is point identified or partially identified. If
a component of $\mathbb{E}[Q_{i}E_{i}]$ is $0$, the upper and lower
bounds for the corresponding component of $\theta$ coincide, and
this component of $\theta$ is point identified. In general, the components
of $\mathbb{E}[Q_{i}E_{i}]$ are not equal to $0$, the bounds do
not coincide, and $\theta$ is only partially identified.
We can assess whether $r_{f}$ is point identified or not by regressing
education $E_{i}$ on friendships $F_{i}$ and other covariates, using
$Z_{i}$ as the instrument, as noted in footnote \ref{fn:bound}.
If the coefficient of friendships is not significantly different from
zero, then we cannot reject the null that $r_{f}$ is point identified
-- the exact returns to education are not important because $F_{i}$
and $E_{i}$ are not correlated. In this case, we would not need to
instrument for education in order to obtain an unbiased estimate of
$r_{f}$. In our data, the coefficient on friendships is significant
from zero -- we can reject the null that $r_{f}$ is point identified.\footnote{This also ensures that the estimated upper and lower bounds are jointly
asymptotically normal, as required by \citet[Assumption 1(i)]{Imbens2004}.}
To be specific about the confidence intervals, denote the upper and
lower bounds of $\theta$ by $\theta_{u}$ and $\theta_{l}$. Let
$\hat{\theta}_{u}$ and $\hat{\theta}_{l}$ be the estimators for
$\theta_{u}$ and $\theta_{l}$ and $se(\hat{\theta}_{u})$ and $se(\hat{\theta}_{l})$
the standard errors of the estimators.\footnote{In Section \ref{sec:age_dist} we also consider an instrument constructed
using estimates from a pairwise regression. In general, standard errors
in two-step estimators should account for the presence of first-step
estimators. However, we are in a special case where the moment condition
in equation (\ref{eq:moment}) has a zero derivative with respect
to the first-step parameters (coefficients in the pairwise regression).
Therefore, the first-step estimation has no impact on the standard
errors in the second step and standard calculation of standard errors
is valid \citep[p.2179]{NEWEY1994}.} \citet[Lemma 4]{Imbens2004} proposed a $(1-\alpha)$ confidence
interval for the true value of $\theta$ that takes the form of $[\hat{\theta}_{l}-c\cdot se(\hat{\theta}_{l}),\hat{\theta}_{u}+c\cdot se(\hat{\theta}_{u})]$,
where the critical value $c$ satisfies $\Phi(c+\frac{\hat{\theta}_{u}-\hat{\theta}_{l}}{\max\{se(\hat{\theta}_{u}),se(\hat{\theta}_{l})\}})-\Phi(-c)=1-\alpha$,
with $\Phi(\cdot)$ being the cdf of the standard normal distribution.\footnote{The confidence intervals proposed by \citet{Imbens2004} require a
superefficient estimator of the length of an identified set (Assumption
1(iii) in their paper). Nevertheless, \citet[Lemma 3]{Stoye2009}
provides a simple sufficient condition for the superefficiency to
hold: the estimated upper and lower bounds are almost surely ordered.
Our estimated bounds are ordered because we take the maximum/minimum
of the two estimates. Therefore, by \citet[Proposition 1]{Stoye2009}
the confidence intervals in \citet{Imbens2004} are appropriate.} In practice, each component of $\theta$ has a critical value $c$,
and we need to solve for it numerically. Using the critical values
for each component of $\theta$, we can construct confidence intervals
for each component of $\theta$.\footnote{The STATA code for the bound estimates and confidence intervals can
be found in Appendix \ref{app:Code}.}
\textbf{Advantage of bounding.} Provided that the true return to education
is between 5 and 15\%, our approach provides consistent bound estimates
and valid confidence intervals for the returns to friendships. This
approach is preferable to a simple calibration that assumes the return
to education is known and equal to a particular value (e.g., $r_{e}=10\%$),
which yields a biased estimator if the true return to education is
different from the presumed value and provides no confidence intervals.
\subsection{\label{sec:age_dist}Using age distance as an instrument}
\citet{McPherson2001} report the results of various studies documenting
that in many settings (including schools) individuals of the same
age are much more likely to be friends. Based on this evidence, we
compute age distance for each pair of individuals as the absolute
difference between their ages to use as an instrument.
The distance between individuals $i$ and $j$ is defined at the pair
level -- that is, between two students of the same school and grade.
To construct an instrumental variable for in-degree at the individual
level, we average the pairwise distances over all $j$ that are in
$i$'s school-grade. We call this measure ``age distance''. In Appendix
\ref{app:pairwise IV} we consider another aggregating approach. We
run a Probit regression of friendships among pairs of students on
the pairwise age distance and basic controls. We then use the predicted
in-degree, calculated by the sum of the predicted friendships, as
an instrument for in-degree.
To be valid our instrument must operate only through friendships and
education (the two endogenous variables). Although age distance also
potentially affects educational attainment (\citealp{black2011too}),
the returns to friendships are still (partially) identified because
we estimate the returns to friendships after subtracting off the impact
of education from log earnings. The instrument satisfies the exclusion
restriction so long as it is uncorrelated with the remaining unobservables.
\subsection{\label{sec:id_var}Identifying variation and identifying assumptions}
We use a similar set of controls and identifying variation as \citet{Murphy2020}
who study the effects of academic rank on academic outcomes, and \citet{Cicala2017}
who show that where an individual ranks within a given social distribution
determines their choice of friends, behaviors, and outcomes. Given
that we control for cohort-mean age and own age, this leaves variation
in age distance among students with the same age and cohort-mean age
for identification.
Table \ref{tab:variation_dist} illustrates this variation for a simple
example. In the first cohort, students are aged 13, 13.5, and 14 (mean
age 13.5). The second cohort has three students aged 13.2, 13.5, and
13.8 (same mean age of 13.5). In both cohorts, there is a student
aged 13.5. However, the age distance for this student is smaller in
cohort 2 (0.3 years) compared to cohort 1 (0.5 years). In the absence
of coordination effects, our model predicts that the student in cohort
2 will have more friends than the student in cohort 1 because they
are closer in age to their peers. The greater variance in the distribution
of ages in cohort 1 compared to cohort 2 results in a greater age
distance in cohort 1 (0.67) than cohort 2 (0.4).
Following \citet{Bifulco2011classmates}, we document in Table \ref{tab:dist_resid}
that there is sufficient variation of this kind in our data after
accounting for the basic set of controls. The variation in age distance,
measured by the standard deviation (s.d. 0.435), is about halved by
the inclusion of individual age (s.d. 0.2). But 45\% of the original
variation (s.d. 0.2) remains after including the full set of controls,
regardless of whether we control for grade FE or mean age in the cohort.
This residual variation in age distance is significant, and larger
than the residual variation in \citet{Bifulco2011classmates}.
A key identifying assumption is that conditional on our basic controls,
age distance does not predict earnings except through its effects
on education and friendships (exclusion restriction). Because the
identifying variation in age distance is generated from differences
in the age distribution across cohorts, we must assume that individuals
do not sort into schools and grades based on the variance in age --
that is, parents do not care or know about age distance and do not
sort using this criterion.
Our identification assumption is not violated if parents want to place
their child in a cohort where the child is older relative to their
classmates (the practice sometimes referred to as red-shirting). In
our previous example, both cohorts have the same mean age, which means
that parents would be indifferent between the two cohorts because
the child\textquoteright s age relative to the cohort mean will be
identical in both cohorts. However, age distance will be on average
smaller in cohort 2. Alternatively, parents may care about their child\textquoteright s
\textit{age rank} within a cohort. In our example, a 13.7 year old
child would still be indifferent between the two cohorts, because
they would be the second oldest in both cohorts. But this student
would be closer in age to their peers in cohort 2, and we expect they
would have more friends.
\subsubsection{\label{sec:exclusion}Empirical Support for the Exclusion Restriction}
While we cannot test the exclusion restriction directly, we conduct
a series of placebo tests using observable pre-determined characteristics
(parental education, parental and own nativity, religion, birthweight
and breastfeeding, height, disabilities, etc.): age distance should
not predict these variables conditional on our basic controls. We
select the variables using two criteria. First, they are mostly determined
early in life, before adolescent friendships are formed. Second, the
prior literature and our data suggest they determine earnings.\footnote{We check that these variables predict earnings by regressing log earnings
with and without the predetermined variables, conditional on the basic
covariates (Columns 1 and 2 of Table \ref{tab:iv_test_joint}). We
reject the null that these variables do not jointly predict earnings
(p-value < 0.001). The $R^{2}$ increases from 0.059 to 0.066 (a 12\%
increase) with the addition of these variables, confirming that they
predict earnings.} Table \ref{tab:iv_test} shows that age distance does not predict
any of the 14 variables we consider, conditional on basic controls:
the coefficients on age distance are statistically insignificant in
all regressions. Moreover, if we regress age distance on all of these
variables (Column 3 of Table \ref{tab:iv_test_joint}), we cannot
reject the null that they do not predict age distance, conditional
on the basic controls (the p-value of the joint F-test is 0.56).
These placebo tests show that many important pre-determined characteristics
that predict earnings are not statistically associated with age distance,
providing support for the exclusion restriction.
\subsection{First stage results: the effect of age distance on friendships}
Table \ref{tab:first_stage} documents that age distance has a negative
and statistically significant effect on in-degree. If the average
age distance between a student and the students in their school-grade
increases by one year, the student has one less friend (Column 1).
Increasing the age distance by one standard deviation (0.44) lowers
the number of friends by about 0.44 friends, a 13\% decline relative
to the mean (3.4). The F-statistic is about 94, well above the standard
thresholds required to rule out weak instruments and large enough
for standard t-statistics in the second stage to be valid \citep{Lee2022}.
These results are similar if we estimate the first stage using the
pairwise data where an observation is a potential link between two
students in the same school-grade (22 million potential links). Using
a Probit probability model, we show that pairwise age distance is
a strong predictor of friendships between two individuals, even after
controlling for the age of the nominated individual in the pair and
the mean age in the cohort (Column 1 of Table \ref{tab:pair_dist}).\footnote{We do not control for the age of the individual who nominates the
friendship because otherwise there is no variation in the pairwise
distance.}
\subsection{Monotonicity}
In settings with heterogeneous treatment effects, IV estimates are
interpretable only under monotonicity: the endogenous variable (number
of friends) must be weakly monotonic in the instrument (age distance)
for all individuals. From a theoretical standpoint, increasing an
individual's homophily level has ambiguous predictions on socializing
and the number of friends because of coordination effects (Proposition
\ref{prop:homophily}). For example, suppose that an adolescent is
placed in a cohort with adolescents that are older instead of being
of the same age. If older individuals socialize more than younger
individuals, then the adolescent may socialize more and accumulate
more friends in the older cohort because it is more productive to
socialize in this group, despite the larger age difference.
Although monotonicity remains an untestable assumption, we empirically
investigate whether it appears to be violated. Figure \ref{fig:first_stage}
shows a bin-scatter plot of friendships and age distance, in the raw
data (left plot) and with basic controls (right plot). The slope is
negative. A non-parametric plot (Figure \ref{fig:first_stage_nonpara})
further confirms that the relationship is weakly decreasing: age distance
does not increase the number of friendships.\footnote{There is a small portion where this is not true but this occurs for
very high values of age distance that are uncommon.} Column 2 of Table \ref{tab:first_stage} shows that age distance
squared has a negative but insignificant effect on friendships. In
the pairwise first stage, age distance squared is significant, so
the relationship is not exactly linear (Column 2 of Table \ref{tab:pair_dist}).
However, the coefficients still imply a monotonic relationship: these
negative coefficients imply that the curve is strictly decreasing
and concave.
We also investigate whether the effect of age distance is symmetric.
Column 3 of Table \ref{tab:first_stage} shows that it is not: it
is more detrimental --- from the point of view of making friends
--- to be young among older peers than it is to be older among young
ones (this is also true in the pairwise estimation, see Column 3 of
Table \ref{tab:pair_dist}). However, age distance is still negatively
correlated to friendships regardless of whether peers are older or
younger.
\citet{Angrist1995} discuss what is needed for the monotonicity assumption
to be met in the case of a continuous treatment. A testable implication
of monotonicity is that the CDFs of the treatment (in-degree) at different
levels of the instrument (age distance) should not cross. A visual
representation of this test is given in Figure \ref{fig:cdf_df}.
We plot the difference in the CDF of in-degree as age distance increases
by one unit.\footnote{Let $X$ denote in-degree and $Z$ age distance. We estimate the CDF
difference $\mathbb{E}[1\{X\leq x\}|Z=z']-\mathbb{E}[1\{X\leq x\}|Z=z]$
for all $x$ and $z'\geq z$ by regressing the indicator variable
$1\{X\leq x\}$ on $Z$. The coefficient of $Z$ provides an estimate
of the CDF difference by one unit increase in $Z$ \citep{Rose2021}.} The CDF differences are non-negative throughout the support, whether
we control for the basic covariates or not. A formal test of this
no-crossing condition is provided by \citet{Barrett2003}. We split
the data evenly into high and low values of age distance. The null
hypothesis is that the distribution of in-degree with high age distance
is first-order stochastically dominated by the distribution of in-degree
with low age distance.\footnote{Let $N_{H}$ and $N_{L}$ denote the number of individuals with age
distance higher and lower than the median in the sample. Let $F_{H}(x)$
and $F_{L}(x)$ denote the CDFs of in-degree for the subgroups with
high and low age distance. The null and alternative hypotheses are
$H_{0}:F_{H}(x)\geq F_{L}(x)$ for all $x$ and $H_{1}:F_{H}(x)<F_{L}(x)$
for some $x$. \citet{Barrett2003} propose the test statistic $\hat{S}=\left(\frac{N_{H}N_{L}}{N_{H}+N_{L}}\right)^{1/2}\sup_{x}(\hat{F}_{L}(x)-\hat{F}_{H}(x))$,
where $\hat{F}_{H}(\cdot)$ and $\hat{F}_{L}(\cdot)$ are estimators
of $F_{H}(\cdot)$ and $F_{L}(\cdot)$. They suggest that the p-value
can be computed by $\exp(-2(\hat{S})^{2})$.} The p-values for the raw and residual distributions are 0.927 and
0.993 respectively, suggesting that we cannot reject the null, further
supporting the monotonicity assumption.
\section{IV Results: The Causal Returns to Friendships}
We now turn our attention to estimating the causal returns to friendships.
\subsection{Reduced form results: homophily and earnings in adulthood}
Figure \ref{fig:earn_iv} shows the correlation between log earnings
and age distance in a bin-scatter plot. Greater age distance is associated
with lower earnings in the raw data (left plot), and with basic controls
(right plot). Conditional on basic controls, individuals in cohorts
with more dissimilar peers in terms of age have lower earnings as
adults: increasing age distance by one standard deviation (0.44) lowers
earnings by 7.5\% (Column 1 of Table \ref{tab:reduced_form}). This
relationship is statistically significant.
\subsection{The causal returns to friendships: IV results}
Table \ref{tab:second_stage_main} reports the estimated bounds on
the returns to adolescent friendships. The OLS estimates for in-degree
are reported in Column 1 for reference, where the return to one more
friend is 0.025. If in-degree is treated as endogenous but education
is not, the estimate for in-degree increases to 0.12 (Column 2). If
we take a calibration approach, where we assume the return to education
is 10\%, the returns to in-degree remain at 0.11 (Column 3).\footnote{This estimate is obtained by running a regression on in-degree, controlling
for covariates and instrumenting for in-degree, where the dependent
variable is given by log earnings minus 10\% times years of schooling.}
Our main specification, which allows for the returns to education
to vary anywhere from 5 to 15\%, bounds the returns to friendships
from 0.093 to 0.137 (Column 4). The confidence interval (CI) for the
returns does not include zero -- these estimates are also statistically
significant. The upper bound corresponds to the lower return to schooling,
and vice versa. This is because the correlation between in-degree
and schooling is positive (Figure \ref{fig:edu_friend} and Table
\ref{tab:cov_edu=000026indeg}): if we regress education on in-degree,
controlling for covariates and instrumenting for in-degree, the coefficient
on in-degree is 0.44 and statistically significant. The significant
coefficient also implies that the returns to friendships are not point
identified.
\textbf{Robustness.} If we use a Probit probability model to predict
links using the pairwise data and then use the predicted in-degree
as an instrument, the bounds are similar and range from 0.065 to 0.096
(Column 5), although the estimates are not significant.
The results are robust to using alternative measures of friendships
(Table \ref{tab:second_stage_measures}), including reciprocated friendships,
which suggests that we are not simply capturing the returns to popularity.
We also check if the results are robust to the inclusion of additional
controls (Table \ref{tab:second_stage_robust}). Column 1 reproduces
our preferred specification for reference and shows bounds of 0.093--0.137.
Our estimates are similar (and the CI does not include 0) if we control
for age rank (Column 2), predetermined individual-level controls (the
ones we used for the placebo tests in Section \ref{sec:exclusion},
Column 3), variables capturing current SES including parental income
and living with parents (Column 4), or their respective cohort-level
means (Column 5). The last two regressions are potentially problematic
because income (or living with parents) is endogenous (it is not truly
predetermined) but it is nevertheless reassuring that the results
are similar.
Table \ref{tab:second_stage_bully} investigates another potential
violation of the exclusion restriction. Perhaps in settings where
age distance is larger, there is greater bullying. Because bullying
can affect mental health \citep{Arseneault2010} which in turn affects
earnings, our instrument could affect earnings through this additional
channel. We test this by controlling for measures of social cohesion
in adolescence. These controls do not individually or jointly affect
the estimated bounds.
\textbf{Magnitudes}. The magnitude of the returns to friendships ranges
from 0.07 to 0.14 across specifications. Since recent estimates of
the returns to schooling are on the larger side (more than 10\%),
the more likely bound for the returns to friendships is the lower
bound, which hovers around 0.07--0.09. Thus an increase of one standard
deviation in the number of friends (3.2) increases earnings by 22\%
-- a large and economically significant return. By comparison, a
one standard deviation increase in years of schooling (2.1) would
increase earnings by 21\% so the effect size is similar.\footnote{The implied elasticity of earnings with respect to friendships is
0.24, whereas the elasticity of earnings with respect to education
is 1.5, assuming the return to education is 10\%.} While the return to in-degree seems large, it represents the returns
to having a friend during adolescence -- we are not measuring the
returns to one year of friendship but the returns to having a friend.
These friendships likely last many years, potentially into adulthood
(Section \ref{sec:discussion}).
Our bounds are larger than the point estimates in \citet{conti2013popularity}
of around 2\%. In addition to methodological differences (they do
not account for the endogeneity of education and friendships), they
study men who were seniors in high school in Wisconsin in 1957, whereas
the Add Health data is nationally representative and surveys students
in middle or high school in 1994--95, at a time when the returns
to education and other individual traits in the labor market are much
larger. They measure labor market outcomes 35 years later whereas
we observe them 15 years later. The returns to friendships in adolescence
may attenuate over time, yielding larger returns when respondents
are surveyed earlier in their careers.
\subsection{Why are OLS estimates downward biased?}
Our preferred bounds do not include the OLS estimate of 0.025.\footnote{However, the confidence interval for the returns to friendships {[}0.008,
0.224{]} includes the confidence interval for the OLS estimator {[}0.018,
0.031{]}, so we cannot reject the null that the IV and OLS estimates
are the same.} We explore two reasons why the estimated bounds are larger than the
OLS estimate: omitted variable bias and measurement error.\footnote{IV estimates might also exceed OLS estimates if treatment effects
are heterogeneous. OLS estimates in Table \ref{tab:subsample} suggest
that while there is some heterogeneity in the returns to friendships
across subpopulations, it is too small to explain the discrepancy
between OLS and IV estimates.}
\textbf{Omitted variable bias.} Omitted variable biases in our model
are in general ambiguous (Proposition \ref{prop:ols}). The sign of
the OLS bias is determined by the correlation between in-degree and
the error term $\epsilon_{i}$.\footnote{Strictly speaking, the bias is given by the covariance between in-degree
and $\epsilon_{i}$, multiplied by the inverse of a matrix $X'X$,
where $X$ consists of residualized education and in-degree. In our
sample this inverse matrix has positive diagonal elements and relatively
small off-diagonal elements (Table \ref{tab:cov_edu=000026indeg}).
Therefore, the sign of the bias is mostly determined by the correlation
between in-degree and $\epsilon_{i}$.} A downward (upward) bias arises if in-degree is positively correlated
with an unobserved factor that negatively (positively) affects earnings.
The data suggest possible omitted factors that could explain our findings:
partying and drinking. Alcohol consumption among adolescents is largely
motivated by its capacity to facilitate social interactions (\citealp{feldman1999alcohol};
\citealp{kuntsche2005young}), but it may have detrimental impacts
on other outcomes. We find that drinking alcohol increases friendships
but lowers GPA and the odds of working in cognitively demanding jobs
(Table \ref{tab:production_functions}).\footnote{We use data from the O{*}NET to construct indices of the extent to
which occupations require social or cognitive skills and match these
scores to a person's occupation. See Appendix \ref{app:ONET} for
details.} Drinking also increases depression and lowers self-reported health
in adulthood. However, time spent with friends is associated with
more friends and better health, lowers the odds of working in jobs
with high cognitive demands, but it does not lower GPA. Thus \textit{certain}
social behaviors, like drinking, increase friendships but lower cognitive
productivity and lower health, both of which likely affect wages.
Additional results suggest that OLS is downward biased. Table \ref{tab:ols_build}
shows that the addition of covariates in OLS regressions of earnings
on grade in-degree increases the estimated returns to friendships.
We control for education only (Column 1), and then we progressively
control endowments (IQ and extroversion - Column 2), personal demographics
(age, gender, race - Column 3), mean characteristics of one's peers
(Column 4), grade fixed effects (Column 5) and school fixed effects
(Column 6). The model with all the controls yields the largest coefficient.
This evidence suggests that individual characteristics and school
fixed effects capture preferences and environments that cause OLS
to be downward biased.
\textbf{Measurement error.} Classical measurement error in friendships
would attenuate the OLS coefficient. The direction and magnitude of
the OLS bias depend on the ratio of the covariance between in-degree
and the true measure, divided by the variance of in-degree: if the
ratio is smaller than one then OLS is attenuated \citep{Bound2001}.
Suppose that high-engagement friendships are the ``true'' number
of friends and grade in-degree is a poor proxy. What would the OLS
bias be in this case? In our sample, the covariance between (residualized)
grade in-degree and (residualized) high-engagement degree is 2.28.
The (residualized) variance of grade in-degree is 8.67 (Table \ref{tab:cov_network_measures}).
Their ratio is about 0.26, which suggests that a consistent estimate
could be 4 times larger than OLS, roughly 0.09, similar to our IV
estimates.
\textbf{Limitations}. Measurement error may also come from the fact
that individuals in the survey are only allowed to list up to five
friends of each gender, generating censoring -- a non-classical form
of measurement error. In this case our IV might be inconsistent --
an IV estimate is consistent only if the instrument is uncorrelated
with measurement error \citep{Bound2001}.
\section{\label{sec:discussion}Discussion: Which Friends Matter and How do
Friendships Affect Earnings?}
We end with an informal discussion of two important issues. First,
do all friendships yield positive returns in the labor market? Second,
what are the mechanisms by which adolescent friendships increase earnings?
While we cannot fully answer these questions in a causal manner, we
discuss some suggestive evidence.
Figure \ref{fig:friend_types} plots the non-parametric relationship
between in-degree and earnings, for different measures of in-degree
suggested by prior studies: same vs. opposite gender (\citealp{McDougall2007};
\citealp{Hall2011}), high vs. low engagement \citep{Gee2017strongties},
high vs. low SES (\citealp{LavySand2019}; \citealp{Fletcher2020};
\citealp{Chetty2022}), and disruptive vs. non-disruptive peers \citep{carrell2018disruptivepeers}.
The friendships that are expected to have higher returns indeed appear
to have higher returns: it pays off more to be friends with people
of the same gender, to have more close friends, to have more friends
who are not disruptive or come from higher SES families. However,
all friendships (except for disruptive ones) appear to have positive
returns, albeit smaller ones. Table \ref{tab:ols_types} confirms
these descriptive patterns using OLS. Unfortunately, our instruments
are not powerful enough to produce precise estimates of different
types of friendships so further work is needed in this area.\footnote{Although all the first stages are strong (Table \ref{tab:first_stage_types}),
we do not have sufficient variation to \textit{separately} identify
the effects of, e.g., male and female friendships.}
Why do friendships in adolescence matter for adult earnings? We do
not have estimates of the causal effect of education on potential
mediators, so we cannot estimate the ideal IV bounds. Instead, we
summarize the suggestive evidence from OLS regressions. Friendships
formed in adolescence are associated with higher likelihood of working
(Column 1 of Table \ref{tab:social_mechanism}), consistent with the
previous literature that friends help workers find jobs (\citealp{Dustmann2015}).
Adolescent friendships are associated with a lower likelihood of working
in repetitive occupations (Column 2) and a greater likelihood of working
in supervisory roles (though this relationship is not statistically
significant, Columns 3). Individuals with more friends are more likely
to be employed in occupations that require greater social skills,
which on average pay larger wages (Column 4). Individuals with more
friends appear to have greater social and management skills and broader
networks as adults. They have more friends in adulthood (Column 5)
and are more likely to be married (Column 6). Controlling for extroversion
in adolescence, those with more adolescent friends are more likely
to be extroverted in adulthood (Column 7). Individuals with more friends
also have greater GPAs (Columns 1 of Table \ref{tab:cognitive_mechanism}),
are less likely to get in trouble while they are in school (Columns
2 and 3), and ultimately are more likely to work in jobs that require
higher cognitive skills (Column 4). They are also less depressed and
healthier (Columns 5 and 6). Thus the returns to friendships in the
labor force do not operate uniquely through social skills and networks
-- friendships also help in the formation of other cognitive and
non-cognitive skills that are rewarded in the labor market.
\section{Conclusion}
We show that individuals that have more friends in adolescence have
higher earnings in adulthood, partly because they are employed in
higher paying occupations that require higher social and cognitive
skills, and partly because they turn into more socially connected
adults that are also more extroverted.
Our results confirm recent findings emphasizing the importance of
adolescence as a formative period for socio-emotional outcomes \citep{Jackson2020}
and particularly for friendships \citep{Denworth2020}. Our findings
also suggest that when students are more similar to each other they
are more likely to form friendships. Therefore, how students are allocated
in groups in schools has important long-term consequences. Re-structuring
classrooms to be more homogeneous in age would appear to benefit all
students involved, from the point of view of friendships and adult
earnings. An important direction for future research is to investigate
which other contexts and student compositions are most conducive to
the creation of friendships.
The importance of socialization during schooling years has further
implications for higher education policy. Many colleges and universities
face criticism for investing in infrastructure for non-academic, recreational
facilities that are believed to contribute to increasing tuition \citep{Jacob2018college}.
These investments are more reasonable in the presence of high returns
to socialization. Our results help rationalize why, for instance,
Greek fraternities and sororities persist, despite the fact that they
can detract from strictly academic endeavors. While partying may decrease
education, particularly if it is associated with drinking, it does
not necessarily decrease earnings. Overall if schools can promote
productive social activities (like working together), it might be
possible to improve both students' social connections and their educational
attainment.
Like recent work (\citealp{Xiang2019}; \citealp{Heckman2012}), our
findings suggest that paying excessive attention to traditional education
measures like test scores might lower the long-term outcomes and well-being
of individuals because they reduce investments in other important
forms of human capital. They also suggest that educational interventions
that are becoming common, such as remote learning, might deter from
social capital formation and result in lost lifetime earnings. In
contrast, interventions that target soft skills might have large returns,
partly through their effects on network formation. Indeed an emerging
literature shows that interventions targeting non-cognitive skills
among young adults can have large returns (\citealp{Katz2022}, \citealp{Heller2017}).
Our paper complements this research by documenting the importance
of having friends, separately from social skills.
There are many unanswered questions that remain. For example, we don't
know much about which environments best foster friendships while maintaining
academic performance. We provide some evidence that not all friendships
matter equally in the labor market but our evidence is only suggestive.
Finally, we only track the impact of adolescent friendships -- not
whether earlier friendships matter, how social networks evolve from
adolescence onward and how adult networks affect labor markets. These
important questions are left to future research.
\bibliographystyle{jpe}
\bibliography{ms.bbl}
\newpage{}
\begin{figure}
\centering\caption{\label{fig:returns_edu_friend}Log Earnings in Adulthood and Friendships
in Adolescence}
\includegraphics{fig_earn_net}\caption*{Note: Add Health restricted-use data. Local polynomial nonparametric plots and 95\% confidence intervals of log earnings on number of friendships in adolescence measured by grade in-degree (see data section for explanations on how grade in-degree is defined). The series separate males and females with non-zero earnings.}
\end{figure}
\newpage{}
\begin{figure}
\centering
\caption{Homophily, Friendships, and Earnings: First Stage and Reduced Form}
\subfloat[\label{fig:first_stage}Friendships and Homophily in Age]{\includegraphics[scale=0.8]{fig_first_stage}}
\subfloat[\label{fig:earn_iv}Log Earnings and Homophily in Age]{
\includegraphics[scale=0.8]{fig_reduced_form}}
\caption*{Note: Add Health restricted-use data. Bin-scatter plots of grade in-degree (panel a) and log earnings (panel b) on age distance with the histogram of age distance (gray bars). In both panels, the left plot uses the original variables, and the right plot uses their residuals after removing individual characteristics (age, IQ, and indicators for whether the student is extrovert, female, and white), cohort-level characteristics (mean age, mean IQ, fraction extrovert, fraction female, and fraction white), grade fixed effects, and school fixed effects.}
\end{figure}
\newpage{}
\begin{table}[t]
\centering \begin{threeparttable}\caption{\label{tab:sum_stats}Summary Statistics}
{\footnotesize{}\begin{tabular}{l*{5}{c}} \toprule &\multicolumn{1}{c}{(1)}&\multicolumn{1}{c}{(2)}&\multicolumn{1}{c}{(3)}&\multicolumn{1}{c}{(4)}&\multicolumn{1}{c}{(5)}\\ &All Students& Extrovert& Shy& Hi IQ& Low IQ\\ \midrule Annual Earnings (W4)& 38073.2& 37753.8& 35312.9& 41988.6& 33578.1\\ & (46674.4)& (43283.4)& (38818.7)& (45419.6)& (45555.9)\\ \addlinespace School In-Degree (W1)& 4.359& 4.733& 3.928& 4.682& 3.998\\ & (3.729)& (3.891)& (3.386)& (3.896)& (3.507)\\ \addlinespace Grade In-Degree (W1)& 3.352& 3.688& 3.069& 3.645& 3.024\\ & (3.153)& (3.347)& (2.847)& (3.271)& (2.975)\\ \addlinespace Grade Out-Degree (W1)& 3.285& 3.484& 3.250& 3.638& 2.883\\ & (2.688)& (2.747)& (2.620)& (2.694)& (2.618)\\ \addlinespace Grade Reciprocated Degree (W1)& 1.388& 1.491& 1.311& 1.597& 1.147\\ & (1.564)& (1.616)& (1.469)& (1.640)& (1.429)\\ \addlinespace Grade Network Size (W1)& 5.248& 5.681& 5.009& 5.686& 4.760\\ & (3.843)& (4.025)& (3.555)& (3.887)& (3.722)\\ \addlinespace Years of Schooling (W4)& 14.76& 14.82& 14.73& 15.37& 14.08\\ & (2.125)& (2.122)& (2.110)& (1.930)& (2.126)\\ \addlinespace GPA (W2) & 2.839& 2.850& 2.818& 2.997& 2.649\\ & (0.746)& (0.739)& (0.758)& (0.735)& (0.714)\\ \addlinespace IQ (W1) & 101.5& 102.1& 100.6& 112.4& 89.30\\ & (13.95)& (13.50)& (14.59)& (7.447)& (9.408)\\ \addlinespace Extrovert (W2) & 0.650& 1& 0& 0.663& 0.632\\ & (0.412)& (0)& (0)& (0.407)& (0.418)\\ \addlinespace Very Hard Study Effort (W1)& 0.381& 0.383& 0.419& 0.346& 0.419\\ & (0.486)& (0.486)& (0.494)& (0.476)& (0.494)\\ \addlinespace Some Study Effort (W1)& 0.511& 0.515& 0.483& 0.536& 0.484\\ & (0.500)& (0.500)& (0.500)& (0.499)& (0.500)\\ \addlinespace Frequently Hang with Friends (W1)& 0.387& 0.401& 0.344& 0.376& 0.399\\ & (0.487)& (0.490)& (0.475)& (0.485)& (0.490)\\ \addlinespace Sometimes Hang with Friends (W1)& 0.523& 0.514& 0.553& 0.541& 0.505\\ & (0.499)& (0.500)& (0.497)& (0.498)& (0.500)\\ \addlinespace Frequently Drink (W1)& 0.163& 0.158& 0.121& 0.174& 0.152\\ & (0.369)& (0.365)& (0.326)& (0.379)& (0.359)\\ \addlinespace Sometimes Drink (W1)& 0.299& 0.306& 0.265& 0.316& 0.283\\ & (0.458)& (0.461)& (0.441)& (0.465)& (0.450)\\ \addlinespace Age (W1) & 16.06& 15.70& 15.83& 16.04& 16.05\\ & (1.670)& (1.534)& (1.546)& (1.619)& (1.716)\\ \addlinespace Age Distance (W1) & 0.980& 0.964& 0.984& 0.903& 1.061\\ & (0.435)& (0.421)& (0.438)& (0.341)& (0.497)\\ \midrule Observations & 10,605& 5,124& 2,765& 5,384& 4,741\\ \bottomrule \end{tabular} }{\footnotesize\par}
\begin{tablenotes}[flushleft]\footnotesize \item
Note: Add Health restricted-use data. Standard deviations in parentheses. W1 stands for Wave I, etc. Due to missing values in IQ and Extrovert, the numbers of observations in Columns 2 and 3 do not sum up to that in Column 1, and the same for Columns 4 and 5.
\end{tablenotes}\end{threeparttable}
\end{table}
\newpage{}
\begin{table}
\centering
\caption{\label{tab:first_stage}What Predicts Friendships in Adolescence?
First-Stage Results}
\begin{threeparttable}{ \def\sym#1{\ifmmode^{#1}\else\(^{#1}\)\fi} \begin{tabular}{l*{3}{c}} \toprule &\multicolumn{1}{c}{(1)}&\multicolumn{1}{c}{(2)}&\multicolumn{1}{c}{(3)}\\ &\multicolumn{1}{c}{In-Degree}&\multicolumn{1}{c}{In-Degree}&\multicolumn{1}{c}{In-Degree}\\ \midrule Age Distance & -1.0760\rlap{***}& -0.7023\rlap{**} & \\ & (0.1111) & (0.2767) & \\ Age Distance Squared& & -0.1120 & \\ & & (0.0721) & \\ Age Distance to Older Peers& & & -2.0300\rlap{***}\\ & & & (0.2822) \\ Age Distance to Young Peers& & & -0.2571 \\ & & & (0.3130) \\ Age & 0.2895\rlap{***}& 0.2534\rlap{***}& -0.5636\rlap{*} \\ & (0.0841) & (0.0813) & (0.3046) \\ Mean Age & -0.3538\rlap{*} & -0.3669\rlap{*} & 0.0784 \\ & (0.1982) & (0.2010) & (0.2314) \\ \midrule F-stat of Homophily Measures& 93.846 & 49.464 & 72.139 \\ P-value & 0.000 & 0.000 & 0.000 \\ $ R^{2}$ & 0.038 & 0.038 & 0.039 \\ Observations & 10,605 & 10,605 & 10,605 \\ \bottomrule \end{tabular} }
\begin{tablenotes}[flushleft]\footnotesize \item
Note: Add Health restricted-use data. In-degree refers to the number of friendships nominated by other students in the same school-grade. Age distance refers to the average age distance to other students in the same school-grade. Age distance to older (younger) peers refers to the average age distance to other students in the same school-grade who are older (younger) than the respondent. Additional covariates include individual characteristics (IQ and indicators for whether the student is extrovert, female, and white), cohort-level characteristics (mean IQ, fraction extrovert, fraction female, and fraction white), grade fixed effects, and school fixed effects. Standard errors in parentheses are clustered at the school level.
\item * p < 0.10, ** p < 0.05, *** p < 0.01.
\end{tablenotes}\end{threeparttable}
\end{table}
\newpage\begin{landscape}
\begin{table}[t]
\centering \begin{threeparttable}\caption{\label{tab:second_stage_main}Labor Market Returns to Friendships:
Main Results}
\begin{tabular}{lccccc}
\toprule
& (1) & (2) & (3) & (4) & (5)\tabularnewline
& OLS & IV & IV & IV & IV\tabularnewline
\midrule
In-Degree & 0.0245 & 0.1242 & 0.1146 & {[}0.0926, 0.1367{]} & {[}0.0650, 0.0956{]}\tabularnewline
95\% CI & {[}0.0177, 0.0312{]} & {[}0.0130, 0.2354{]} & {[}0.0155, 0.2138{]} & {[}0.0075, 0.2242{]} & {[}-0.0308, 0.1936{]}\tabularnewline
\midrule
First Stage for In-Degree & & Aggregate & Aggregate & Aggregate & Pairwise Probit\tabularnewline
Education Endogenous & N & N & Y & Y & Y\tabularnewline
Education Returns & 10.30\% & 7.83\% & 10\% & 5-15\% & 5-15\%\tabularnewline
Observations & 10,605 & 10,605 & 10,605 & 10,605 & 10,605\tabularnewline
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]\footnotesize \item
Note: Add Health restricted-use data. The dependent variable is the log annual earnings. In-degree refers to the number of friendships nominated by other students in the same school-grade. Column 1 presents OLS estimates for the returns to in-degree. Column 2 presents estimates where in-degree is endogenous, but education is not (estimated returns to education are reported in the second to last row). Column 3 presents estimates where in-degree is endogenous and the returns to education is assumed to be 10\%. Columns 4-5 present estimates where in-degree is endogenous and the returns to education are bounded between 5 and 15\%. All specifications include individual characteristics (age, IQ, and indicators for whether the student is extrovert, female, and white), cohort-level characteristics (mean age, mean IQ, fraction extrovert, fraction female, and fraction white), grade fixed effects, and school fixed effects. The confidence intervals are constructed based on standard errors clustered at the school level.
\end{tablenotes} \end{threeparttable}
\end{table}
\newpage{}
\begin{table}[t]
\centering \begin{threeparttable}\caption{\label{tab:second_stage_robust}Labor Market Returns to Friendships:
Robustness Checks}
\begin{tabular}{lccccc}
\toprule
& (1) & (2) & (3) & (4) & (5)\tabularnewline
& IV & IV & IV & IV & IV\tabularnewline
\midrule
In-Degree & {[}0.0926, 0.1367{]} & {[}0.0908, 0.1347{]} & {[}0.0952, 0.1356{]} & {[}0.0951, 0.1355{]} & {[}0.0694, 0.1157{]}\tabularnewline
95\% CI & {[}0.0075, 0.2242{]} & {[}0.0068, 0.2212{]} & {[}0.0002, 0.2333{]} & {[}0.0030, 0.2303{]} & {[}0.0033, 0.1841{]}\tabularnewline
\midrule
Additional Individual Controls & N & Age Rank & Pre Controls & SES Controls & N\tabularnewline
Additional Cohort Mean Controls & N & N & N & N & SES Controls\tabularnewline
Observations & 10,605 & 10,605 & 10,605 & 10,605 & 10,605\tabularnewline
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]\footnotesize \item
Note: Add Health restricted-use data. The dependent variable is the log annual earnings. In-degree refers to the number of friendships nominated by other students in the same school-grade. All specifications include individual characteristics (age, IQ, and indicators for whether the student is extrovert, female, and white), cohort-level characteristics (mean age, mean IQ, fraction extrovert, fraction female, and fraction white), grade fixed effects, and school fixed effects. Column 2 controls for age rank, which refers to the ranking of the respondent's age in the school-grade. Column 3 controls for additional individual-level predetermined covariates, including number of siblings, mother's age at the student's birth, birth weight, height, and indicators for whether the mother was born in US, the father was born in US, the parents are Catholic, Baptist, the student was born in US, was breastfed, is mentally retarded, and has diabilibity. Column 4 controls for additional individual-level SES controls, including family income, mother's years of schooling, father's years of schooling, and indicators for whether the mother lives in the household and whether the father lives in the household. Column 5 controls for cohort-level means of the additional SES controls in Column 4. The instrument is age distance. The confidence intervals are constructed based on standard errors clustered at the school level.
\end{tablenotes} \end{threeparttable}
\end{table}
\end{landscape}
\pagestyle{appendix}\newpage{}