Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Engineering Social Networks: How Initial Group Assignment Shapes Student Social Interactions
abstractA large literature uses exogenous variation to estimate how assignment to classrooms or other groups shapes social networks. Yet most of these analyses remain dyadic, treating each link in isolation, even though ties often form through triadic closure, as a friend of a friend also becomes a friend. Using fine-grained data on phone calls, text messages, physical co-location, and social-media ties, we estimate the network formation effects of randomly assigning first-year university students to classrooms and to smaller social groups. To analyze explicitly whether group assignment interact with triadic closure, we use our random assignment to estimate a subgraph generated model of network formation. Accounting for triadic closure turns out to be crucial. For social groups in particular, group assignment affects network formation almost entirely by inducing additional triadic closure. Estimates ignoring triadic closure can thus yield misleading predictions about the network effects and benefits of group assignment policies.
Code and synthetic data are available at: \url{https://github.com/johanna-einsiedler/engineering-social-networks}
bibunit[ecta-fullname]
\section{Introduction}
Our lives are profoundly shaped by the social networks we are embedded in, which channel access to opportunities, information, and social capital and thereby affect behavior at both the individual and aggregate level jackson_social_2008,easley_networks_2010.
Existing work finds causal social influence of peers and social links on fundamental aspects of human life, including academic achievement booij_ability_2017, epple_chapter_2011, mehta_time-use_2019, social behavior duncan_peer_2005,bursztyn_social_2017, and occupational choice marmaros_peer_2002. As a result, policymakers often seek to intervene in these network formation processes to further various policy goals.
Existing causal studies evaluate how organized groups affect whether two people connect, but by treating each potential tie in isolation they say nothing on how groups affect higher-order structure - in particular triadic closure, where shared connections draw people together - even though that structure is what makes networks cohesive.
Social links do not form at random. Researchers have identified three mechanisms behind who connects with whom rivera2010dynamics.
First, social foci---such as communal activities, non-private physical spaces, and social events---provide the opportunity for interaction and relationship building feld_focused_1981.
Second, homophily describes the tendency of individuals to associate with those who are similar in various characteristics mcpherson_birds_2001, which in turn shapes social integration and positive spillovers golub2012homophily,bramoulle2012homophily,bjerre2020assortative.
Third, triadic closure posits that two individuals are more likely to form a link if they share at least one mutual connection (thereby forming a triangle) Watts1998,kossinets_empirical_2006,
which plays an essential role in fostering cooperation and trust as well as coordinating behavior jackson2012social, breza2019social.
In practice, many policy interventions targeting networks operate by reshaping social foci within people's private, professional or academic lives. Examples of such interventions are the random assignment process of college students to dorm rooms (e.g. sacerdote_peer_2001, sacerdote_peer_2001) or the allocation of high school students to classrooms (e.g. graham_teacher--classroom_2020, graham_teacher--classroom_2020). To design social-foci interventions effectively, it is necessary to understand not only how changes in social foci causally affect network formation directly, but how it propagates through triadic closure - since the same intervention can build scattered ties or knit dense clusters depending on which channel dominates.
In this paper, we estimate the causal effect of two policy interventions that aim to change social networks in higher education by modifying social foci. When arriving for their first semester at the Technical University of Denmark (DTU), we randomly assigned students within each study program to different academic classrooms for their teaching assistant (TA) sessions, with approximately 30 students in each classroom. Independently, we also randomly assigned the same students to social groups of around seven within each program, formed to welcome them to university. Crucially, we do not only estimate the overall effect of each intervention on the propensity to connect. We separate that effect into two channels: the direct effect of a shared focus on whether a given pair links, and its effect through triadic closure, whereby a shared focus creates the mutual connections that pull further pairs together. To our knowledge, no prior study identifies these two channels separately.
To estimate these effects of interest, we overcome several key challenges that pervade the literature on network formation. First, to separate causal effects from endogenous selection, we leverage the random assignment of the 2013 cohort, which we implemented with DTU's administration and student organizations. Because late arrivals and non-random regrouping meant assignment did not perfectly determine final membership, we measure students' realized groups and instrument them with their randomly assigned groups. Second, to reliably measure social connections, we use rich smartphone-based data from the Copenhagen Network Study sapiezynski_interaction_2019. This allows us to measure students' actual social interactions via phone calls, messages, Facebook friendship and time spent together physically.
Third, separating a direct link from one induced by triadic closure is not just a measurement problem but an identification problem: the two channels are bundled in any observed network. We address this by embedding a subgraph generated model (SUGM) of network formation by chandrasekhar_network_res in our randomized design through a control function.
This lets us recover the causal effect of shared group membership on direct linking and on triadic closure separately - turning the SUGM, which describes higher-order formation, into a tool for causal identification of it.
We start our analysis by estimating the overall causal effect of changing social foci on students' social interactions in a standard dyadic framework. Our random group assignment procedure generates exogenous variation in whether a pair of students is assigned to the same social group or classroom. We use this assignment to instrument realized shared membership in a two-stage least squares (2SLS) framework. We find strong effects of social groups on interactions. Students assigned to the same social group in their first semester meet 2.9 more times per week in their second semester than students who were not assigned to the same group. Similarly, students in the same social group talk on the phone for 17 more seconds per week, send 1.3 more text messages, and are 4.4 times as likely to be Facebook friends. Moreover, looking at time horizons beyond the second semester reveals that most of the effects are declining but persistent. The effect of classrooms is weaker and limited to fewer dimensions. Students in the same classroom in their first semester have an average weekly call duration that is about two seconds longer, with the effect being significant at the 5% level. By contrast, the effects on Facebook friendship, weekly SMS count (0.11 higher), and weekly physical meetings are positive but small and not statistically significant at the 10% level.
Then, we apply principal component analysis to combine data on physical meetings, phone calls, and text messages into a single binary measure of which students have formed links with each other. Based on this, we estimate that two students placed in the same social group are 32 percentage points more likely to form a link (on a baseline of 9%). For classrooms, the effect is 3 percentage points (on a baseline of 8%). These estimates capture the total effect on linking. They do not, on their own, reveal whether shared foci build ties directly or by inducing closure - the decomposition we turn to next.
Next, we examine triadic closure effects. First, we adapt the SUGM framework of network formation to leverage random assignment for identification in a similar way as in the dyadic 2SLS framework. In a SUGM, a link can arise two ways. It may form directly between a pair of individuals as a function of their characteristics and circumstances. Or it may be the byproduct of triadic closure when two individuals each link to a third individual. As noted by chandrasekhar_network_res, the latter implies that the model is able to capture triadic closure effects, while remaining parsimonious and tractable. We adapt their SUGM model to allow for the standard identification issues around non-random sorting into groups and classrooms and then show how our random group assignment can be used to estimate causal effects using an instrumental variable and a control function approach. In the adapted model, shared membership can affect both whether a pair links directly and whether any triad they belong to closes into a triangle. Results show that social group effects indeed interact strongly with triadic closure. Being in the same social group substantially increases the likelihood that three students form a triangle. The estimated effect on pairwise linking formation, by contrast, is small, statistically insignificant and if anything negative. Together these suggest that shared group membership shapes the networks primarily by drawing clusters of students together rather than strengthening pairwise connectivity - though the direct channel is estimated imprecisely enough that we cannot rule out a modest positive effect. For classrooms, results are not precise enough to reliably disentangle pairwise link and triadic closure effects, however, estimates point primarily to effects on pairwise link formation with triadic closure effects playing a limited role.
Finally, we use the estimated dyadic and SUGMs to simulate counterfactual networks ($N=500$ draws per model) and to quantify how accounting for triadic closure changes substantive conclusions. The simulations show that the two models can produce similar link density while implying very different higher-order structure: relative to the dyadic specification, the SUGM generates substantially more triangles and higher clustering. We then translate these structural differences into welfare differences using a parsimonious utility framework in which students benefit from direct links and from spillovers through mutual friends. When spillover strength is small, the models yield similar welfare predictions; as spillovers become more important, the SUGM predicts higher average utility.
Taken together, the simulation exercise shows that triadic closure is not only a statistical feature of formation but also a substantively consequential mechanism for evaluating interventions that shift social foci. In particular, policies that change classroom or social-group composition can have effects that operate primarily through the creation (or suppression) of clustered friendship structures---and these effects will be understated if researchers rely on models and estimates that ignore triadic closure.
Overall our paper relates and contributes to a large literature on network formation and peer effects. Previous work has documented the importance of social foci---such as spatial proximity and organizational grouping---in driving the formation of social ties. For instance, marmaros2006friendships show that dormitory proximity plays a pivotal role in establishing friendships among college students, while hallinan1985ability provide evidence that ability grouping influences student interactions. These studies underscore that initial group composition is crucial for the subsequent structure of social networks. Recent advances in measurement have further enriched our understanding of network formation by overcoming systematic error in survey data, caused by e.g. social desirability bias or incomplete knowledge. Mobile phone data, for example, have been used to infer friendship network structure reliably onnela2007structure, eagle2009inferring, stopczynski2014measuring, stadtfeld2015partnership. These studies illustrate that modern data sources can capture the dynamics and link strengths within networks, thereby providing a more nuanced view of how early interactions emerge and evolve.
Our approach contributes to the literature on measuring the impact of social foci on how social networks form. We leverage random assignment to groups which allows us to isolate the causal effect of initial social contact on formation of group structures. This design is similar in spirit to the methodology employed by sacerdote_peer_2001, who exploit random roommate assignments to uncover peer effects, and to the structural framework advanced by griffith2024random for assessing counterfactual treatment effects under non-random peer influences. The main novelty relative to this work is that we study not only the overall propensity to form links, but also the impact on triadic closure.
Methodologically, our paper contributes by showing how subgraph-based network formation models can be combined with random assignment (or other exogenous variation) to credibly establish causal effects. This enables us to synthesize diverse strands of literature---from the micro-foundations of network ties mcpherson_birds_2001, feld_focused_1981 to the causal identification of peer effects sacerdote_peer_2001, kremer_peer_2003---and demonstrates that early, exogenously induced groups have lasting effects on the structure of social networks.
Beyond immediate interactions, a growing literature has emphasized that early peer contacts yield persistent spillovers on academic achievement and socio-economic trajectories. For example, goette2012impact document how minimal group settings can influence collective behavior, while altmejd2021brother and barrios2022neighbors demonstrate that siblings and neighbors socially influence college major choices across countries. Our findings resonate with these studies by showing that the peer networks forged through random assignment not only affect initial interactions but also have lasting consequences for the structure of social networks.
Our work further points to the importance of non-linear peer effects stemming from the subgroup formation within organizationally assigned groups.
\section{Data and setting}
Our analysis focuses on the social networks of new undergraduate students at the Technical University of Denmark (DTU). DTU offers various undergraduate and graduate educational programs in civil engineering. First-year students at DTU enter one of several bachelor programs or diploma programs, ranging from Biomedical Engineering, to Cyber Technology and Mechanical Engineering.
Upon joining DTU, new students are assigned to two key social foci by the university administration and student organizations: first, within each study program, DTU administration splits students into two or more teaching assistant classroom groups containing around 30 students each. All teaching assistant (TA) sessions take place within these assigned groups. Second, in collaboration with the administration, an organization of older students groups new students into introductory social groups, each consisting of about 7 students from the same study program. Throughout the first semester, these groups meet with a senior “group mentor” and conduct various social introduction activities such as e.g. an inaugural weekend trip. Participation in this social group system is voluntary, however, a majority of students sign up to join a social group.\footnote{Out of the 1,372 students that we have data on, we observe 1,069 joining a social group during the first semester.}
Both classroom and introductory social group assignments at DTU are prime examples of the types of social foci typically leveraged for policy interventions. By manipulating assignment of different students to these groups, policy makers may be able to shape the social networks that students form. The aim of our analysis is to understand the causal effects of such group assignments on the properties of the students' social networks.
\subsection{Random assignment in the 2013 cohort}
A key concern when estimating the causal effects of assignment to classrooms or social groups is the possibility of endogenous selection. Even when assignment is formally done by administrators or student organizations, students with certain unobservable characteristics can often be more (or less) likely to end up in the same group. The direct effect of such unobservables can bias the estimated effects of group assignment.
To overcome this challenge, our analysis exploits random assignment to groups. In mid-August 2013, a few weeks prior to the start of the semester, we assisted the administration and student groups at DTU in randomly assigning incoming students to TA classrooms and introductory social groups. In both cases the randomization of students was performed conditional on gender to comply with existing rules about the potential gender mix within groups.\footnote{Specifically, the administration and student organization impose certain restriction on the minimum number of female students that should be assigned to a given group or classroom.} For the social groups, randomization was additionally conditioned on whether the student had dietary restrictions, as this was requested by the student organization.\footnote{To ease planning of the various social activities, students with dietary restrictions are always grouped together when forming groups.} We account for this conditioning in our analysis (see Section (ref)). Balance tests across socio-demographic characteristics show no significant differences between groups (see Supplemental Material, Section A), supporting successful randomization.
Importantly, the final classroom and social group assignment used in practice may differ from our assigned groups for two reasons: First, our randomization was applied to the list of students who had signed up with the DTU administration and/or requested membership in a social group by mid-August. Late sign-ups were thus not part of our randomization. Second, the student organization specifically reserved the right to make minor adjustments to the social groups if the need arose. We account for both of these sources of imperfect compliance in our analysis.
\subsection{Administrative data}
The basic data set for our analysis was obtained from the university administration and covers students enrolled in the 2013 cohort. In addition to information about which study program the student is enrolled in, this data includes basic demographic information, such as sex and age. We supplement this data with information on the randomized classroom and social group assignment discussed above. Further, we obtained updated data from the university administration in 2024, containing information on the actual classroom and social group that each student was placed in during the academic year.\footnote{The administrative data on classrooms was reported in a different format than the original classroom assignment and contained more detailed information on all courses a student has been enrolled in during their first semester. We thus took a conservative approach and marked every pair of students as being members of the same classroom that shared at least one course. We provide a robustness check using a different definition in the Supplemental Material (Section E).}
\subsection{Measuring networks and social interactions: The Copenhagen Network Study}
To construct a comprehensive and reliable measure of networks, we leverage a unique data set on social interactions generated by the Copenhagen Network Study (CNS) stopczynski2014measuring,sapiezynski_interaction_2019. Over 700 students studying at DTU in 2013 participated in the study. Participants of this study were handed out Android-based Google Nexus 4 smartphones and asked to install data collection software from Google Play Store. The experiment was GDPR compliant and registered with the Danish authorities (for a more comprehensive discussion and detailed information on the data collection process see sapiezynski_interaction_2019).
We use three kinds of social interaction data collected as part of CNS:
Bluetooth data. The smartphones were set to request and receive responses from all nearby Bluetooth discoverable devices every five minutes. These responses included unique identifiers of devices within a maximum range of 10 meters that were connected to the study participants. For each pair of users, we counted a maximum of one meeting per 15-minute interval. In cases where Bluetooth sensors detected users in close proximity multiple times within the same 15-minute period, only the first instance was recorded as a meeting. Further, the location of each meeting was recorded and subsequently categorized as being on or off the DTU campus. Additionally, based on the administrative class scheduling data obtained from DTU, we computed indicators for individual class attendance, based on whether individuals are physically present in the classrooms bjerre2020negative.
Calls and short messages. Metadata on calls and short messages (SMS) were obtained from the smartphones on a daily basis. Participants were required to reveal their phone number and could thus be matched to the call logs. For the case of SMS, each record consists of the unique identification numbers of the sender and receiver as well as a timestamp, specifying the time of the event. Similarly for calls, identification numbers of the person calling and the person receiving the call were logged, the timestamp when the call started and the call duration in seconds.
Facebook data. In addition, the majority of students voluntarily opted in to allow data collection from Facebook. A snapshot of the students friendship network was collected at the end of the observation period.
It should be noted that the data were collected in 2013, when Facebook was considerably more popular as a social medium and SMS was used far more frequently as a primary communication channel than today.
\subsection{Samples}
Our analysis relies on two different main samples, one aimed at estimating the effects of classroom assignment and one aimed at estimating the effects of social groups. We describe their construction below, while Table (ref) provides an overview of sample totals.
\subsubsection{Social Groups Sample}
As per the official data obtained from the DTU administration in mid-August, 1,372 students enrolled in the 2013 DTU undergraduate cohort. Since fewer students had signed up for a social group by mid-August 2013, however, restricting to students that were part of our social group randomization cuts the sample to 616. To analyze the effect of actual group membership, we restrict attention to 606 students who we also see in the final social group assignment data from fall semester 2013. To be able to construct data on social interactions and networks, we also restrict to students who participated in the Copenhagen Network Study ($n=223$). Finally, since our analysis studies group assignment and network formation within each study program, we restrict attention to study programs where we observe at least 10 students. This leaves us with a final sample of 171 students for use in analyzing social groups.
\subsubsection{Classroom Sample}
The construction of our analysis sample for studying classrooms proceeds analogously to above. However, here we have more complete data, with random assignment information being available for the full sample and actual classroom membership information available for 901 students. After imposing all the other restrictions described above, we are left with a sample of 276 students which we use to analyze the effects of classrooms.
\begin{table}[ht]
\caption{Overview of the sample totals for the social group sample and the classroom sample}
\begin{tabularx}{\linewidth}{lYYY}
\hline
& Social Group Sample& Classroom Sample \\
\hline
Initial group assignment information available & 616 & 1,372 \\
Corrected group assignment information available & 606& 1,143 \\
Data on sex available & 578 & 1,103\\
Participation in Copenhagen Network Study & 223 &409\\
Part of a study major with min. 10 observations & 171 &392\\
\hline
\end{tabularx}
\end{table}
\subsection{Measuring social linking}
\subsubsection{Dyadic data structure and social outcomes}
For our main analysis, we adopt a dyadic data structure; individual observations are pairs of students within the same study program and we aim to examine whether being in the same social group or classroom affects social interactions between the two students constituting the pair.
Accordingly, from each of our two student-level analysis samples we form all possible pairs, $ij$, where students $i$ and $j$ are from the same study program $s$.\footnote{The two social foci we analyze, classrooms and social groups, are only ever shared by students from the same study program. Comparisons across students from different study programs would thus never be meaningful in our analysis.} In each of the resulting dyadic data sets, we further construct the variables $SameGroup_{ijs}$ as a dummy variables for whether $i$ and $j$ are in the same social group or classroom, respectively. Analogously, we construct the variable $SameAssignedGroup_{ijs}$ as a dummy for whether the pair initially had been assigned to be in the same social group/classroom in our randomization. As noted earlier, $SameGroup_{ijs}$ and $SameAssignedGroup_{ijs}$ may differ because of imperfect compliance.
Our main outcomes of interest are the social interactions and network formation, as measured based on calls, SMS, physical co-presence and Facebook friendship.
For easier interpretation, we compute weekly averages for each of the four semesters in our data\footnote{\textbf{Semester 1 - Fall Semester 2013:} 1\textsuperscript{st} September 2013 - 31\textsuperscript{st} of January 2014; \textbf{Semester 2 - Spring Semester 2014:} 1\textsuperscript{st} February 2014- 31\textsuperscript{st} of July 2014; \textbf{Semester 3 - Fall Semester 2014:} 1\textsuperscript{st} August 2014- 31\textsuperscript{st} of January 2015; \textbf{Semester 4 - Spring Semester 2015:} 1\textsuperscript{st} February 2015- 31\textsuperscript{st} of July 2015 }. For each continuous outcome variable, this is done by summing all observed pairwise interactions and dividing by the number of weeks in the semester.
We construct the following variables for all pairs with the relevant data available: (1) average weekly call duration in minutes, (2) average weekly SMS count, (3) average weekly meetings outside the DTU campus and (4) a binary indicator for whether a given pair became friends on Facebook before the end of the observation period or not.
In our main analysis, we focus on interactions in students' second semester at DTU, i.e. the Spring semester of 2014. We do this to avoid social interactions that happen mechanically due to scheduled activities in the social group and/or around the actual TA sessions taking place in the first semester. In the Supplemental Material (Section ), we examine effects on other semesters to see whether the effects on interactions persist at longer time horizons.
\subsubsection{Composite measure of social links}
In addition to estimating the effect of group membership on different types of social interactions, we are also interested in providing an overall measure of the effect on networks and link formation. To achieve this, we construct a composite measure that captures social interaction across multiple channels and establish a criterion for determining whether a given pair has formed a link.
To do this, we first use principal component analysis (PCA) to aggregate the three continuous measures of pairwise social interactions into a single aggregate continuous measure. We then use this aggregate to define which students are linked.
For our main analysis, we code a pair of students as linked if their score on the aggregate interaction measure is above the 90\textsuperscript{th} percentile within their study program.
Table (ref) provides a summary of the resulting network structures across study programs. Since we only consider pairs within the same study program, each study program effectively constitutes its own network. Across the 20 study programs where we have all necessary information for at least 10 individuals, we observe average degrees between 2.3 to 6.9, with a median of 3.9. This indicates that our cutoff aligns well with the notion of a social circle of having a “handful” of friends within the study program. The maximum of the observed diameters is 6, meaning that two individuals within the same study program are at maximum separated by 6 other individuals. The minimum average path length is 1.3, compared to a maximum of 2.7, showing that some networks are considerably more dispersed than others. It should be noted that, because the dataset only includes students who participated in the study, interactions with individuals outside the sample are not observed. Consequently, the resulting network represents only a partial view of the students' social environment and network statistics might be subject to bias KOSSINETS2006247.
We provide further detail on the construction of the link measure, detailed statistics for every study program and robustness checks for this in the Supplemental Material (Section B).
\begin{table}
\caption{Summary of network topology measures across all 20 study programs based on the composite social links measure and a cutoff at the 90\textsuperscript{th} percentile within in each study program. Each study program constitutes its own network as we only look at pairs within the same study program.}
\begin{threeparttable}[t]
\begin{tabularx}{\linewidth}{lYYYYY}
\hline
& Min. & 25\textsuperscript{th} percentile & Median & 75\textsuperscript{th} percentile & Max.\\
\hline
Number of pairs & 66 &136 & 200 & 253 & 595 \\
Number of links & 7 & 14 & 20 & 26 & 60\\
Average degree & 2.33 & 3.29 & 3.90& 4.52 & 6.86 \\
Clustering coefficient & 0.00 & 0.31 & 0.41 & 0.50& 0.80\\
Diameter & 2 & 3 & 4&5&6\\
Average path length &1.30&1.71&2.26&2.36&2.74\\
\hline
$n=20$\\
\hline
\end{tabularx}
\end{threeparttable}
\end{table}
\section{Overall effects of group assignment on interactions and link formation}
We start by estimating the overall effect of shared membership in a classroom or social group using a simple dyadic specification. We do this in the context of the following simple linear regression, applied to our dyadic samples of student pairs from the same study program:
\begin{equation}
Y_{ijs} = \beta_0 + \beta_1 SameGroup_{ijs} + v_{ijs}
\end{equation}
The outcome $Y_{ijs}$ here is some measure of pairwise social interactions and/or link formation between students $i$ and $j$. $SameGroup_{ijs}$ is either our dummy variable for being in the same TA classroom or the same social group. The coefficient of interest is $\beta_1$ which measures the causal effect of shared group membership on the outcome variable. Note that, as written, the model assumes that the causal effect is homogeneous. If targeting a Local Average Treatment Effect (LATE) interpretation, however, our main instrumental variable approach can allow for heterogeneous treatment effects in the standard way ImbensAngrist1994.
Potential endogenous selection into groups means that simple estimation of (ref) using ordinary least squares (OLS) is likely to give a biased estimate of $\beta_1$. Instead we rely on the variation generated by our random assignment in an instrumental variables setup. Recall that $SameAssignedGroup_{ijs}$ is a dummy for whether $i$ and $j$ were assigned to the same group in our randomization. Under the realistic assumption that our assignment influenced the actual groups formed, $SameAssignedGroup_{ijs}$ should be a relevant instrument for $SameGroup_{ijs}$ in (ref).
To be a valid instrument, $SameAssignedGroup_{ijs}$ also needs to satisfy the necessary exclusion restriction, i.e. being assigned to the same group must not affect interactions other than through the actual group membership.
Despite stemming from a randomization, there are two concerns: First, randomization was done conditional on some observed characteristics. Second, our randomization did not directly randomize the dummy $SameAssignedGroup_{ijs}$ but randomized all students in a given study program into a number of different groups. Thus, while on an individual level group assignment is random conditional on covariates, as pointed out forcefully by borusyak_nonrandom_2023, such a randomization need not guarantee exogeneity of $SameAssignedGroup_{ijs}$ in (ref).\footnote{A very concrete problem in our setting is the size differences across study program. In a study program with many students, a given pair of students is systematically less likely to end up in the same group under our randomization; this makes $SameAssignedGroup_{ijs}$ systematically related to the size of $i$ and $j$'s study line. At the same time, study lines may well differ in the extent of social interactions and network formation.}
A convenient way to address both concerns is to use the “instrument recentering” approach proposed by borusyak_nonrandom_2023. Intuitively, this means adjusting the observed value of the instrument to account for the baseline probability of two individuals being assigned to the same group by chance. For a more formal explanation see Supplemental Material (Section C).
To implement this practically, we re-run our random assignment procedure 1,000 times and compute the share of realizations where each pair $ijs$ is assigned to the same group (see Section C of the Supplemental Material for the exact randomization algorithm and distribution of group assignment probabilities). We then subtract this average from $SameAssignedGroup_{ijs}$ to arrive at a \emph{recentered} instrumental variable which we denote by $\widetilde{SameAssignedGroup_{ijs}}$. We use this as our instrument in a two-stage least squares (2SLS) estimation of (ref), with the following first stage:
\begin{equation}
SameGroup_{ijs} = \gamma_0 + \gamma_1 \widetilde{SameAssignedGroup_{ijs}} + u_{ijs}
\end{equation}
Estimating linear dyadic specifications like the ones above using 2SLS is common in the literature, see e.g. Algan2023, harmon_peer_2019. Briefly foreshadowing later results, we note that a numerically equivalent way to estimate $\beta_1$ using random assignment here would be to use a control function approach and include estimated residuals from (ref) as controls in (ref). We adopt such an approach later when estimating our SUGM model.
For inference, we need to address the concern that errors terms are likely to be correlated across pairs of students who share a member. To account for this, we use dyad robust standard errors based on samii_cluster-robust_2015 for our main specification. Supplemental Material (Section E) presents alternative results using bootstrapped standard errors based on resampling study programs. Results are almost identical to those obtained using dyad robust standard errors. \footnote{Dyadic robust errors may fail to account for all dependencies in the data under second order network effects and/or triadic closure effects. Additionally, the bootstrap results ensure direct comparability with the triadic closure results presented later which use the same bootstrap approach for inference.} As an additional robustness check we also estimate our main specification with study program fixed-effects and a different definition of TA classroom membership. For the fixed-effects results change only slight, the alternative classroom definition results in a smaller sample and results in similar effect size with slightly higher standard errors (see Supplemental Material, Section E).
\subsection{Effect on measures of link formation}
Figure (ref) shows the results for 2SLS estimates of the effects of group membership on interactions, based on the recentered instrument. The leftmost column of (teal) bars shows results for social groups, while the rightmost column of (red) bars shows results for classrooms. Within each row and column, the first bar shows the predicted \emph{baseline} levels of social interactions among pairs not in the same group ($\beta_0$), while the second shows the predicted levels of social interaction among pairs who are part of the \emph{same group} ($\beta_0+\beta_1$). Error bars reflect the standards error on the difference between baseline and same group pairs (the standard error on $\beta_1$). For completeness, Table (ref) reports numerical estimates. In the Supplemental Material (Section D) we additionally report the 2SLS specifications reduced form and first stage estimates.\footnote{Unsurprisingly, for both social groups and classrooms, the first stage relationship between shared group membership and (recentered) assigned shared members is strong. For social groups, the first stage coefficient on the instrument is around 0.76. For classrooms it is around 0.96. Accordingly, as shown in Table (ref), the usual first-stage F-statistics for instrument strength are large.}
For social group membership we find significant effects across all forms of interactions. Over 90% of students in our sample that were in the same social group eventually became Facebook friends, compared to a baseline of 21% of students in the same study program but not the same social group. For SMS and call duration we find particularly large effects compared to the baseline. On average, members of the same social groups spend 19 seconds on the phone with each other every week, compared to the baseline of 1 second. We also find that members of the same social group text each other more; on average they exchange 1.4 SMS per week compared to only 0.08 at baseline. Finally, students in the same social group are also much more likely to meet physically outside of campus, 4.7 times per week as opposed to only 1.8 times for baseline pairs not in the same group.\footnote{When it comes to physical interactions, we restrict our main analysis to interactions taking place outside of the university campus, as we are mainly interested in the effect of shared group membership on interactions not enforced through study related activities (e.g. group work), we chose to exclude any meetings happening on the university campus for our main specification. However, one could easily argue that a significant part of friendship-induced meetings could also happen on campus (e.g. going for lunch in between classes, studying together). We thus report results for (1) the average number of weekly meetings outside class hours and (2) the overall average number of weekly meetings (i.e. including during classes) in the Supplemental Material (Section F). Effect sizes are larger for these alternative specifications and exhibit identical significance patterns. The effects reported here can thus be seen as a conservative lower bound of the effect of group membership on physical meetings.}
\begin{table}
\begin{threeparttable}[t]
\caption{Group membership effects on social interactions: Results from 2SLS estimation based on re-centered group assignment instrument}
\begin{tabularx}{\linewidth}{lYYYYYY}
\hline
Outcome variable & Facebook & Weekly call duration & Weekly SMS count & Weekly physical meetings & Binary friendship indicator\\
& (1) & (2) & (3) &(4) &(5)\\
\hline
\multicolumn{3}{l}{\emph{Panel A. Social Groups}} &&\\
Same group & 0.696***&0.290***&1.334**&2.928*** & 0.317***\\
&(0.063)&(0.096)&(0.529)&(1.052) & (0.067)\\[.5cm]
Intercept & 0.207***&0.022**&0.083**&1.775*** & 0.085***\\
&(0.029)&(0.010)&(0.035)&(0.304) & (0.015)\\ [.5cm]
Dyads &1,002 & 1,057 & 1,057 & 1,082 & 1,057\\
Students &156 & 160 & 160 & 162 & 160\\[.5cm]
1\textsuperscript{st} stage $F$& 1,311.75 & 1,539.29 & 1,539.29&1,511.12& 1,539.29\\
\hline
\multicolumn{3}{l}{\emph{Panel B. Classrooms}} \\
Same group & 0.028 & 0.035** & 0.110 &0.079 & 0.026*\\
&(0.030)& (0.017)&(0.069)&(0.209) & (0.015)\\ [.5cm]
Intercept & 0.274***&0.035***&0.101***&1.637*** & 0.083***\\
&(0.027)&(0.011)&(0.038)&(0.222) & (0.014)\\[.5cm]
Dyads & 3,468 & 3,459 & 3,459 & 3,459 & 3,491\\
Students & 362 & 260 & 260 & 262 & 364\\ [.5cm]
1\textsuperscript{st} stage $F$ & 2,667.30 &2,848.18&2,848.18&2,864.81,2,841.31\\
\hline
\end{tabularx}
\begin{tablenotes}[para]
• \textit{Notes:
This table shows instrumental variable estimates of the effect of being in the same group on various interaction outcomes during the second semester, i.e. spring term 2014. The baseline mean
refers to the average level of interactions among pairs who don't share a group. The coefficients represent the effect of being in the same group. Group membership is instrumented by group assignment. Dyad robust standard errors are in parentheses. }\\
• \textit{ * p $<$ .1, ** p $<$ .05, *** p $<$ .01.}
\end{tablenotes}
\end{threeparttable}
\end{table}
\begin{figure}[h!]
\caption{Average weekly interactions for pairs not sharing a social group vs. pairs sharing a social group (left panel) as well as for pairs not sharing a classroom vs. pairs sharing a classroom (right panel), based on 2SLS with recentered instrument. }
\end{figure}
For classrooms, the effects are less pronounced. We neither find an effect of classroom membership on the likelihood of becoming Facebook friends nor physical meetings off campus. We do, however, observe comparatively small but significant effects on call duration and SMS count. Pairs within a classroom spend an average of 6 seconds on the phone with each other per week (compared to 3 seconds for those not sharing a classroom) and exchange $\sim1$ SMS every month (compared to a baseline of 0.6).
Finally, in the Supplemental Material (Section F), we show that the estimated effects appear persistent. When looking at social interactions in later semesters we find that effect sizes do decline but for phone interactions the estimates remain significant.
\subsection{Effect on overall link formation}
In addition to measuring the overall effect of group membership on social interactions, we are also interested in providing a summary measure of the effects on link formation. The last column of Table (ref) shows the estimated effects of social group and classroom membership on our summary measure (using a cutoff at the 90\textsuperscript{th} percentile) of links. For social groups, being in the same group increases the likelihood of forming a link by 32 percentage points, corresponding to a 3.8 fold increase compared to the baseline of 8%. For classrooms, the effect is less pronounced. While the effect is nonetheless significant, it is only 3 percentage points, thus roughly increasing the baseline likelihood by a third.\footnote{For robustness and transparency, in the Supplemental Material (Section B) we report alternative results using different versions of our aggregate measure of links. Overall, effect sizes are similar.}
\section{Interactions between group assignment and triadic closure}
We finish our empirical analysis by examining possible interactions between our social foci group interventions and triadic closure. Doing so requires an extension to our empirical specification. Being based on dyadic analysis of pairs, the specifications in the previous section are well-suited to study the overall effects on link formation. By focusing only on pairs, however, the specifications are unable to account for triadic closure effects which inherently operate across three students simultaneously.
Empirical models of triadic closure effects can raise challenges of complexity and computational cost. To mitigate this, we rely on recent work by chandrasekhar_network_res showing that so-called Subgraph Generated Models (SUGM) are able to fit triadic closure patterns well with a relatively small number of parameters and without incurring excessive computational cost. Specifically, we show how the same random-assignment variation used in the dyadic 2SLS framework can be to estimate the SUGM in an instrumental-variable control-function approach.
\subsection{A Subgraph Generated Model}
A simple way to understand SUGMs is that they model network formation and measurement \emph{as if} happening through a two-step process: In the first (unobserved) step, a number of different types of graphs form independently between individuals. In the simplest possible form with triadic closure, these subgraphs are just pairwise links and triangles consisting of links across three individuals. For each pair of individuals there is thus some probability that a pairwise link is formed. Separately, for each triad of individuals there is also some probability that a triangle of links form between them. These probabilities typically depend on the characteristics of the pair or triad, including whether they have any shared social foci such as being in the same social group or TA classroom.
In the second step, the final network is observed as the \emph{union} of the pairwise links that have formed \emph{plus} all links that are part of some triangle that has formed. In the observed data, a link between a given pair of individuals may thus reflect that a pairwise link formed in the first step and/or that some triangle involving this pair formed in the first step.
\subsubsection{Effect of groups on link and triangle formation}
We now formally describe the SUGM that we estimate in our analysis. As in the previous sections, we analyze social groups and classrooms separately, e.g. we set up and estimate a separate SUGM for studying social groups and classrooms.\footnote{When analyzing the effect of social groups, the effect of classrooms will thus be subsumed in the flexible unobservable terms included in the model, and vice versa when analyzing the effects of classrooms.} In what follows below, the term “group” will thus refer to either social groups or classrooms.
First, in terms of data and notation, we extend our analysis data to not only contain all within-study program pairs $ij$ but also all within-study program triads of students $ijk$. Correspondingly, we let $Triangle_{ijks}$ be a dummy variable for whether students $i$, $j$ and $k$ from study line $s$ are all linked to each other based on our summary measure of links between students. Letting $\Omega_s$ denote all the students in study line $s$, we also introduce notation for the sets of all unique student pairs and student triads in the study line, $\Omega_s^L=\{(i,j)\in \Omega_s^2:i>j\}$ and $\Omega_s^T=\{(i,j,k)\in \Omega_s^3:i>j \,\wedge j>k \}$.
Second, to explicate the two-step structure of the SUGM, we let $L_{ijs}^{direct}$ be a dummy variable for whether student $i$ and $j$ from study program $s$ formed a pairwise link in the first step of the SUGM. Similarly, let $T_{ijks}^{direct}$ be a dummy variable for whether the triad $ijk$ formed a triangle in the first step of the SUGM. Note that these variables are fundamentally unobserved. Additionally, as shorthand notation let $\boldsymbol{G}$ be a vector summarizing the group assignment of all students in the data.
In the first step of the SUGM, the likelihood that a given pair $ij$ forms a link or that a triad $ijk$ forms a triangle will depend on group memberships as well as a number of other factors and characteristics, which we summarize by the (unobserved) scalars, $\eta_{ijs}$ and $\eta_{ijks}$, respectively. We let $\boldsymbol{\eta}$ denote the vector of all such unobserved scalars across all pairs and triads in the data. We normalize these unobservables to be mean zero.\footnote{As we explicate below, the unobservables will enter into regression equations with intercepts ($\beta_0$ and $\theta_0$). Assuming that the error terms are mean zero thus simply normalizes the intercepts to a specific and easily interpretable value}
As a natural extension of the linear regression framework of the previous sections, we now assume that the likelihood of forming a pairwise link depends linearly on whether the pair is in the same group and on the scalar unobservable:
\begin{align}
E\left[L_{ijs}^{direct}|\boldsymbol{G},\boldsymbol{\eta}\right]=\beta_0 + \beta_1 SameGroup_{ijs} + \eta_{ijs}
\end{align}
As before, $\beta_1$ is the first key coefficient of interest, representing the causal effect of $i$ and $j$ being in the same group on the likelihood of forming a link in the first step of the SUGM. Similarly, for the likelihood that a triad forms a triangle, we assume that it depends linearly on the number of pairs within the triad who are in the same group plus the scalar summarizing unobservables:
\begin{align}
E\left[T_{ijks}^{direct}|\boldsymbol{G},\boldsymbol{\eta}\right]=\theta_0 + \theta_1 \sum_{\left(\iota,\kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} } SameGroup_{\iota \kappa s} + \eta_{ijks}
\end{align}
The key coefficient of interest here is $\theta_1$ which is defined as the causal effect of having one more pair within the triad be in the same group. Switching from a situation where all three members in the triad are in different groups to a situation where two of the students in the triad are in the same group thus increases the likelihood of forming a triangle by $\theta_1$. Switching to a situation where they are all in the same group increases the likelihood of forming a triangle by $3 \theta_1$.\footnote{Note that mechanically, the number of pairs within a triad that are in the same group is either 0, 1 or 3. It is not possible for two pairs of students to be in the same group without having all three students in the triad in the same group.}
In order to impose that the model has the SUGM structure, we further impose that conditional on group memberships and scalar unobservables, the likelihood of forming a link is independent of the likelihood of forming a triangle that is
\begin{align*}
\left\{ L_{ijs}^{direct} \right\}_{(i,j)\in \Omega_s^L} \cup \left\{ T_{ijks}^{direct} \right\}_{(i,j,k)\in \Omega_s^T} \quad \textrm{ mutually independent given } \left(\boldsymbol{G},\boldsymbol{\eta}\right)
\end{align*}
Note that at this point, the assumptions above are just a normalization. Since we have imposed no restriction on the dependence of the scalar unobservables $\boldsymbol{\eta}$ across the different pairs or triads of students, the assumptions above can fit any unconditional dependence patterns in the unconditional likelihood of forming links and triangles. Assumptions further below will however impose additional restrictions.
In terms of the observed data, the final network is the union of the links and triangles formed in the first step. This means that in our observed network there will be a link between $i$ and $j$ if such a pairwise link formed in the first step \emph{or} if some triangle formed that contains both $i$ and $j$:
\begin{align}
Link_{ijs}=
\begin{cases}
1, & \text{if} \quad L_{ijs}^{direct}=1 \quad \textrm{or} \quad T_{ijks}^{direct}=1 \textrm{ for some } k \\
0, & \text{otherwise}
\end{cases}
\end{align}
Finally, to see that the SUGM model outlined above is a natural extension of the dyadic 2SLS framework of the preceding section, note that if we set $\theta_0,\theta_1,\eta_{ijks}=0$ (e.g. remove the possibility of forming triangles in the first step), then the SUGM collapses exactly to the linear regression (ref) from Section (ref) with $\beta_1$ and $\beta_0$ having the same values and interpretation.
If $\theta_0$, $\theta_1$ and $\eta_{ijks}$ are not all zero however, the model explicitly includes triadic closure effects and if $\theta_1>0$ in particular, these effects interact with shared group membership. The model then implies that the final network will contain a relatively large share of triangles of three interconnected students, compared to overall number of links between students. Moreover this pattern will be exacerbated among triads of students who are placed in the same group. Estimating $\theta_1$ will thus allow us to determine whether our social foci interventions on groups interact with triadic closure.
\subsubsection{Identification and estimation using random assignment}
Analogously to Section (ref), the fundamental challenge for identifying the causal effects of shared group membership, is that we typically would expect shared group membership to be systematically related to many potential other factors that affect link and triangle formation, e.g. $SameGroup_{ijs}$ will not be independent of $\eta_{ijs}$ and $\eta_{ijks}$.
As before, we solve this using the variation generated from our random group assignment. Since a SUGM with triadic closure inherently becomes non-linear, we integrate the instrumental variables approach using control functions. Our instrument will again be the recentered dummy for whether given pair of students have been assigned to the same group, $\widetilde{SameAssignedGroup_{ijs}}$. The instrument first stage is assumed to be as before:
\begin{equation}
SameGroup_{ijs} = \gamma_0 + \gamma_1 \widetilde{SameAssignedGroup_{ijs}} + u_{ijs}
\end{equation}
As a matter of notation, we let $\boldsymbol{Z}$ and $\boldsymbol{u}$ denote the vectors stacking the value of the instrument and the first stage error, $u_{ijs}$, across all pairs of students.
To accommodate the SUGM structure, we then proceed to impose slightly stronger versions of the standard instrumental variables assumptions. First, we restrict the first stage relationship between the instrument and actual group membership. Just as above, we assume that the coefficient on the instrument in the first stage is non-zero and that conditional on the value of the instrument, the error term in the first stage is mean zero:
\begin{align}
\gamma_1 \ne 0 \\
E\left[u_{ijs}|Z_{ijs}\right]=0
\end{align}
Next we invoke assumptions analogous to the exclusion restriction used previously but now strengthened to match the SUGM and allow for a control function approach.
Let $\xi_{ijs}=\eta_{ijs}-E\left[\eta_{ijs}|\boldsymbol{u},\boldsymbol{Z}\right]$ and $\xi_{ijks}=\eta_{ijks}-E\left[\eta_{ijks}|\boldsymbol{u},\boldsymbol{Z}\right]$ denote the deviations of the unobservables from their mean conditional on the instruments and first stage errors. To permit a standard control function approach we next assume that this conditional mean does not depend on the instrument but only on the first stage errors corresponding to the involved pairs of students:
\begin{align}
E\left[\eta_{ijs}|\boldsymbol{Z},\boldsymbol{u}\right]=h_{L}(u_{ijs}) \\
E\left[\eta_{ijks}|\boldsymbol{Z},\boldsymbol{u}\right]=h_{T}(u_{ijs},u_{iks,}u_{jks})
\end{align}
The interpretation of these assumptions are that the first stage (ref) decomposes the variation in same group membership into a part stemming from the random assignment instrument, the term involving $\widetilde{SameAssignedGroup_{ijs}}$ and a part stemming from endogenous selection of certain pairs of students into the same group, $u_{ijs}$. The assumptions above then correspond to saying that the latter of these contains information about the mean of the endogenous link and triangle formation unobservables, $\eta_{ijs},\eta_{ijks}$ but the former does not.
Next, to preserve the SUGM structure within the control function framework, we impose that the deviations of the unobservables from their conditional mean are mutually independent conditional on the first stage errors and instruments:
\begin{align}
&\left\{ \xi_{ijs}\right\} _{(i,j)\in \Omega_s^L} \cup \left\{ \xi_{ijks}\right\} _{(i,j,k)\in \Omega_s^T} \quad \textrm{ mutually independent given } \left(\boldsymbol{Z},\boldsymbol{u}\right)
\end{align}
The key feature of SUGM models that make them so empirically tractable is that they split link and triangle formation into processes that are independent. The assumption above ensures that this continues to hold also once we condition on the instruments and first stage errors.
\subsection{Estimation using control functions and non-linear least squares}
Together the assumptions above allow us to estimate the key causal parameters of interest, $\beta_1$ and $\theta_1$, using a two-step control function approach explained in detail in the appendix.
To conduct inference, we apply a (block) bootstrap approach and resample study programs 100 times. This corresponds to an assumption that the formation of links and triangles is independent across study lines but requires no additional restrictions on dependence within a study program. We show the distribution of the resulting effect sizes across draws in the Supplemental Material (Section H).
\subsection{Results}
Table (ref) shows estimates from the SUGM outlined above both for the effects of social groups and classrooms. For social groups, we see a significant effect of shared group membership on the likelihood of forming a triangle (in the first step of the SUGM). For a given triad of students, having two of them be in the same social group increases the likelihood of forming a triangle by 3.7 percentage points. If instead all three are in the same group, the likelihood increases by a total of 11.1 percentage points. Conversely, for a triad of students where none are in the same group, the estimate is in fact slightly negative, suggesting virtually no chance that a triangle forms. These results confirm that social group membership interacts very markedly with triadic closure. Indeed, looking at the effects of social group membership on link formation in the first step, the estimate is negative and insignificant. In other words, results suggest that the overall effects of social groups on link formation is driven primarily by the formation of triangles.
Turning to the results for classrooms, estimates are more noisy. For pairwise link formation, the estimated effect of classroom membership are sizable, suggesting that being in the same classroom positively impacts the likelihood that a given pair forms a link (in the first step of the SUGM). Conversely, the estimates suggest that classroom membership has no impact on the likelihood of forming a triangle. Taken together these results suggest that a classrooms affect network formation very differently than social groups, with classroom effects operating exclusively at the pairwise-level. At the same time, standard errors are large enough that none of the effects are significant. At conventional level of significance we are thus not able to draw any firm conclusions.
For our main specification we use a quadratic control function and give equal weights to both dyads and triangles in the optimization procedure. As explained in the appendix, another natural choice would be to weight triangles proportional to the ratio of dyads to triangles. We implement this specification a robustness check and find that effects are all in the same direction, of similar magnitude and at the same significance level as our main specification. Additionally, we also estimate the model with linear or cubic
control functions. This yields qualitatively similar results, with a positive and significant effect of shared group membership on triangle formation and no robust effect on direct link formation. All robustness checks are reported in the Supplemental Material (Section H).
\begin{table}
\begin{threeparttable}[t]
\caption{Control-function SUGM estimates of link and triad formation}
\begin{tabularx}{\linewidth}{lYYYY}
\hline
& \multicolumn{2}{c}{Social groups} & \multicolumn{2}{c}{Classrooms} \\
& Pairs & Triads & Pairs & Triads\\
& (1) &(2)&(3)&(4)\\
\hline
Same group ($\beta_1$ or $\theta_1$)& -0.027 & 0.037***& 0.120 & -0.005 \\
& (0.110)&(0.007) & (0.117)&(0.005)\\
Intercept ($\beta_0$ or $\theta_0$) & 0.031 & -0.004* & -0.007 & 0.012\\
&(0.036)&(0.003)&(0.078)&(0.010)\\
\hline
\end{tabularx}
\begin{tablenotes}[para]
• \textit{Notes:
This table shows the estimates of the effect of being in the same group on link and triad formation, based on nonlinear least-squares estimation of the parameters in the SUGM. The baseline mean
refers to the likelihood of a given pair (triad) forming a link (triangle) when none of the individuals share a group. The coefficient in the link equation represents the effect of a pair sharing a group on their likelihood of forming a link. The coefficient in the triangle equation represents the effect of a triad forming a triangle if one additional individual shares a group with any of the other members of the triad. Standard errors are based on 100 bootstrap samples.}\\
• \textit{ * p $<$ .1, ** p $<$ .05, *** p $<$ .01.}
\end{tablenotes}
\end{threeparttable}
\end{table}
\section{Simulations and welfare implications}
\subsection{Utility framework}
To interpret the importance of modeling triadic closure, we evaluate simulated networks through a welfare/utility lens in which individuals value (i) direct links and
(ii) spillovers that arise from being indirectly connected through mutual acquaintances. Let
$A$ denote the (undirected) adjacency matrix of the network, with $A_{ij}=1$ if $i$ and $j$
are linked and $A_{ij}=0$ otherwise, and let $d_i=\sum_{j\neq i}A_{ij}$ be node $i$'s degree.
We normalize the utility from a direct link to be $U_L$ per endpoint, so that direct-link
utility is
\begin{equation}
U^{(i)}_{\text{direct}}(A)=U_L\sum_{j\neq i} A_{ij}=U_L\cdot d_i .
\end{equation}
Spillovers depend on whether $i$ and $j$ have at least one common neighbor, i.e. whether
$(A^2)_{ij}>0$. We distinguish spillovers from \emph{open triads} (a friend-of-a-friend
relationship without a direct link, $A_{ij}=0$) and \emph{closed triads} (a direct link that is
embedded in at least one triangle, $A_{ij}=1$). Individual spillover utility is
\begin{equation}
U^{(i)}_{\text{indirect}}(A)
=
U_L \sum_{j\neq i}
\Big(
\alpha\,\mathbbm{1}\{(A^2)_{ij}>0\}\,(1-A_{ij})
+
\beta\,\mathbbm{1}\{(A^2)_{ij}>0\}\,A_{ij}
\Big),
\end{equation}
where $\alpha\in(0,1)$ captures the value of open-triad spillovers and $\beta>0$ captures the
additional value of closed-triad embeddedness. For convenience we also use the re-parameterization
$\gamma=\alpha+\beta$ (overall magnitude of spillovers) and a mixing parameter $\lambda$
that governs the relative weight placed on open- versus closed-triad spillovers (equivalently,
$\alpha$ and $\beta$ can be expressed as functions of $(\gamma,\lambda)$). Total utility is then
\begin{equation}
U_i(A)=U^{(i)}_{\text{direct}}(A)+U^{(i)}_{\text{indirect}}(A),
\qquad
\bar U(A)=\frac{1}{n}\sum_{i=1}^n U_i(A).
\end{equation}
This framework makes clear that if $\gamma=0$ (no spillovers), welfare depends only on the
number of links; if $\gamma>0$, welfare also depends on higher-order structure, in particular
the prevalence of triangles and friend-of-friend relationships.
\subsection{Simulation design}
To evaluate the implications of the different network estimation frameworks under the proposed utility model, we simulate (i) a purely dyadic model (link formation driven only by pair-level terms), and
(ii) the SUGM specification that allows links to form either directly at the pair level or
indirectly as part of triangle formation (triadic closure). As our unconstrained estimation of the SUGM yields partially negative coefficients, we re-estimate the same model with the additional constraints that the intercepts and the sum of intercept and same group effects need to be positive. The resulting estimates are of the same sign. For social groups they only differ slightly in magnitude and show the same significance levels. For classrooms we find that under the constraints, the effect sizes are smaller and significant. They are reported in the Supplemental Material (Section H).
For each model, we generate
$N=500$ simulated networks using the estimated parameters and the observed grouping
structure (study line membership and social group membership). Figure (ref)
illustrates typical draws from each model. Below we comment on how they differ in terms of network structure and its consequences for welfare.
\begin{figure}[h]
\begin{subfigure}{0.42\textwidth}
\caption{ Example of a simulated network from the dyadic model}
\end{subfigure}
\begin{subfigure}{0.45\textwidth}
\caption{ Example of a simulated network from the SUGM model}
\end{subfigure}
\caption{Examples of simulated networks. Color indicates studyline membership, shade indicates social group membership.}
\end{figure}
\subsection{Results}
\begin{table}[t]
\caption{Number of observed links and triangles averaged across simulations based on the dyadic model and the SUGM model. N=500.}
\begin{tabular}{l|cc}
\hline
& Avg. links & Avg. triangles \\
\hline
Dyadic Model & 130 & 9.52 \\
SUGM & 133 & 51.6\\
\hline
\end{tabular}
\end{table}
A first takeaway is that the two models generate networks with \emph{similar link density}
but \emph{very different clustering}. Table (ref) shows that the average
number of links is nearly identical across the two models (about 130 vs.\ 133), whereas the
SUGM produces far more triangles.
Thus, a dyadic model can match first-order moments (links) while substantially
under-predicting the higher-order structure that is central to triadic closure.
We next translate these structural differences into welfare differences by computing $\bar U(A)$
for each simulated network and comparing the dyadic and SUGM predictions. Figure (ref)
plots the mean difference in utility (SUGM minus dyadic) across simulations as a function of
spillover strength $\gamma$ and the spillover-mix parameter $\lambda$.
When spillovers are small (low $\gamma$), the welfare gap is close to zero: because both models generate similar numbers of links, they imply similar utility when
utility is dominated by direct connections. As spillovers become more important (higher
$\gamma$), the welfare gap increases sharply, reflecting that the SUGM generates substantially
more triangles and thus more opportunities for indirect benefits and embedded relationships.
Moreover, the gap varies systematically with $\lambda$, consistent with the fact that changing the
relative importance of open- versus closed-triad spillovers changes how much the additional
triadic structure produced by the SUGM is valued.
Overall, the simulations show that allowing for triadic closure can substantially change counterfactual welfare evaluations. In settings where policy objectives depend on clustering, cohesion, or spillovers through mutual
friends, dyadic models thus risk understating the welfare consequences of interventions that operate
through group structure and cliques.
\begin{figure}[h]
\caption{Mean utility difference across $N=500$ simulated networks (SUGM minus dyadic), shown over the spillover strength $\gamma$ and the mixing parameter $\lambda$ that governs the relative weight on open- versus closed-triad spillovers.}
\end{figure}
\section{Conclusion}
This paper evaluates how group membership---in small groups with a social purpose and larger groups with an academic purpose---affects interactions between first-year students at university.
To achieve this, we first exploit initial random group assignment and estimate a dyadic 2SLS model. Further we introduce a novel empirical strategy, estimating a subgraph generated model using random assignment. We use this model to simultaneously estimate causal effects on link and triad formation.
Our findings contribute to a better understanding of multiple aspects of the driving forces of social network formation. First, we study how group membership in a university setting affects actual interactions between students. We find large, significant and persistent effects for social groups with around 7 members each and small effects of mixed significance for classrooms with a size of around 30 students each. Based on these results, it becomes clear that assignment to social groups constitutes a potent intervention for engineering friendship circles, at least under the conditions studied by us.
Second, we find evidence that social groups primarily influence interactions through the formation of small cliques (i.e., triads) rather than just pairwise connections. In contrast, our estimates for classrooms are noisier but suggest no significant effect on triangle formation, indicating that network formation processes operate very differently across group types.
Third, our simulation results show that incorporating triadic closure is not merely a modeling refinement but can meaningfully change welfare and policy conclusions. Networks simulated from the dyadic and SUGM specifications can exhibit similar link density yet markedly different clustering, with the SUGM generating substantially more triangles. When we evaluate these counterfactual networks using a simple utility framework that allows for spillovers through mutual friends, the implied welfare differences grow with the strength of spillovers. This highlights that interventions such as social-group assignment may improve outcomes primarily by fostering clustered friendship structures and the associated indirect benefits, effects that would be understated by models that ignore triangle formation.
Overall, our results suggest that assigning students to small groups with a primarily social purpose can be an effective way to shape the structure of friendship networks. At the same time, several limitations should be considered when interpreting these findings. First, participation in the study is voluntary, raising the possibility of self-selection into the sample. Students who choose to participate may differ systematically from non-participants, for example in their sociability or engagement with university life, which could affect the observed interaction patterns. Second, the external validity of our results may be limited by the specific institutional context of the study. The data come from a relatively homogeneous cohort of first-year students at a Danish technical university in 2013, and the relatively low levels of homophily we observe may partly reflect this setting. Third, our empirical approach relies on independence assumptions inherent to the SUGM framework, which abstract from potentially richer forms of dependence in network formation beyond links and triangles.
These considerations caution against extrapolating the quantitative magnitudes of our estimates to other contexts. Nevertheless, the randomized assignment of students to groups provides a rare opportunity to study causal drivers of network formation in a natural institutional environment. By combining this design with a framework that explicitly incorporates triadic closure, our analysis highlights how institutional structures can shape not only the prevalence of social ties but also the clustered architecture of friendship networks. More broadly, our results underscore the importance of accounting for higher-order network structure when evaluating interventions that aim to influence social interactions.
\section*{Appendix - Details of control function estimator of the SUGM model}
This section provides the derivations underlying the control function estimator for the SUGM model.
In addition to the setup and assumptions from the main text, we use the following additional notation: We define $L_{ijs}^{incidental}$ as a dummy variables for whether a triangle has formed in the first step which implies that there is a link between $i$ and $j$ from study program $s$ in the final network:
\begin{align}
L_{ijs}^{incidental}=
\begin{cases}
1, & \text{if} \quad T_{ijks}^{direct}=1 \textrm{ for some } k \\
0, & \text{otherwise}
\end{cases}
\end{align}
Below we will refer to this also as the pairwise link forming \emph{incidentally}.
Similarly,we define $T_{ijks}^{incidental}$ as a dummy for whether a triangle has formed incidentally between $i$, $j$ and $k$ from study line $s$. This occurs if each of the links in the triangle has either formed directly or has formed as part of some other triangle:
\begin{align}
T_{ijks}^{incidental}=
\begin{cases}
1, & \text{if} \quad \forall\,\left(\iota \kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\}:\, L_{\iota \kappa s}^{direct}=1\vee\max_{\mu\in\varOmega_{s}\setminus\{i,j,k\}}T_{\iota \kappa \mu s}^{direct}=1 \\
0, & \text{otherwise}
\end{cases}
\end{align}
Finally, to shorten notation we let $Z_{ijs}$ denote the (recentered) instrument and $G_{ijs}$ denote the pairwise same group dummy:
$$Z_{ijs}=\widetilde{SameAssignedGroup_{ijs}}$$
$$G_{ijs}=SameGroup_{ijs}$$
and also introduce shorthand notation for the conditional probabilities that links and triangles form directly:
\begin{align}
p_{ijs}^L=P\left(L_{ijs}^{direct}=1|\boldsymbol{G},\boldsymbol{\eta}\right) \\
p_{ijks}^T=P\left( T_{ijks}^{direct}=1|\boldsymbol{G},\boldsymbol{\eta}\right)
\end{align}
\subsection{Conditional probability of forming a link or triangle in the final network}
We start by deriving expressions for the conditional probability of forming (seeing) pairwise links and triangles in the final observed network.
First we note that the (conditional) probability that a pairwise link
forms incidentally is equal to the probability that at least one triangle
containing the pair forms directly. Working step by step we can then
write:
\begin{align*}
P\left(L_{ijs}^{incidental}=1|\boldsymbol{G},\boldsymbol{\eta}\right) & =P\left(\max_{k\in\varOmega_{s}\setminus\left\{ i,j\right\} }T_{ijks}^{direct}=1|\boldsymbol{G},\boldsymbol{\eta}\right)\\
& =1-P\left(\max_{k\in\varOmega_{s}\setminus\left\{ i,j\right\} }T_{ijks}^{direct}=0|\boldsymbol{G},\boldsymbol{\eta}\right)\\
& =1-\underset{k\in\varOmega_{s}\setminus\left\{ i,j\right\} }{\Pi}P\left(T_{ijks}^{direct}=0|\boldsymbol{G},\boldsymbol{\eta}\right)\\
& =1-\underset{k\in\varOmega_{s}\setminus\left\{ i,j\right\} }{\Pi}\left(1-p_{ijks}^T\right)
\end{align*}
The second line simply writes things in terms of the probability that
none of the involved triangles form. The third line uses the independence underlying the SUGM model, (ref).
With this we can characterize the (conditional) likelihood of observing
a given link in the data. It happens for sure if the link forms directly,
but even if the link does not form directly, it could form incidentally:
\begin{align}
P\left(L_{ijs}=1|\boldsymbol{G},\boldsymbol{\eta}\right)= & p_{ijs}^L+\left(1-p_{ijs}^L\right)P\left(L_{ijs}^{incidental}=1|\boldsymbol{G},\boldsymbol{\eta}\right)\nonumber \\
= & p^L_{ijs}+\left(1-p^L_{ijs}\right)\left(1-\underset{k\in\varOmega_{s}\setminus\left\{ i,j\right\} }{\Pi}\left(1-p^T_{ijks}\right)\right)
\end{align}
Next we go through similar steps for the conditional probability of observing
a triangle in the data. A triangle forming incidentally means that
each of its three links must form directly or as part of some other
triangle forming directly. Whether these links form and/or these other
triangles form however are independent events.\footnote{Note a key fact here is that each other triangle can overlap with
at most one link in the triangle under consideration (because two
triangles either share 0, 1 or 3 links).} Going through steps similar to the ones used just above gives us:
\begin{align*}
P\Big( & T_{ijks}^{incidental}=1|\boldsymbol{G},\boldsymbol{\eta}\Big)=\\
&P\left(\forall\,\left(\iota \kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\}:\, L_{\iota \kappa s}^{direct}=1\vee\max_{\mu\in\varOmega_{s}\setminus\{i,j,k\}}T_{\iota \kappa \mu s}^{direct}=1 |\boldsymbol{G},\boldsymbol{\eta}\right)\\
& =\underset{\left(\iota \kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }{\Pi}P\left(L_{\iota \kappa s}^{direct}=1\vee\max_{\mu\in\varOmega_{s}\setminus\{i,j,k\}}T_{\iota \kappa \mu s}^{direct}=1|\boldsymbol{G},\boldsymbol{\eta}\right)\\
& =\underset{\left(\iota \kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }{\Pi}\left(1-P\left(L_{\iota \kappa s}^{direct}=0\wedge\max_{\mu\in\varOmega_{s}\setminus\{i,j,k\}}T_{\iota \kappa \mu s}^{direct}=0|\boldsymbol{G},\boldsymbol{\eta}\right)\right)\\
& =\underset{\left(\iota \kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }{\Pi}\left(1-P\left(L_{\iota \kappa s}^{direct}=0|\boldsymbol{G},\boldsymbol{\eta}\right)\underset{\mu\in\varOmega_{s}\setminus\{i,j,k\}}{\Pi}P\left(T_{\iota \kappa \mu s}^{direct}=0|\boldsymbol{G},\boldsymbol{\eta}\right)\right)\\
& =\underset{\left(\iota \kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }{\Pi}\left(1-\left(1-p^L_{\iota \kappa s}\right)\underset{\mu\in\varOmega_{s}\setminus\{i,j,k\}}{\Pi}\left(1-p^T_{\iota \kappa \mu s}\right)\right)
\end{align*}
The second and fourth line uses the independence in (ref).
The other lines simply write out complementary probabilities.
With this expression can then write the probability of observing a
given triangle by noting that it will be observed if formed directly
but otherwise could still be observed because it formed incidentally:
\begin{align}
P\!\left(T_{ijks}=1 \mid \boldsymbol{G},\boldsymbol{\eta}\right)
=& p^T_{ijks} + \Bigl(1-p^T_{ijks}\Bigr)
P\!\left(T_{ijks}^{incidental}=1 \mid \boldsymbol{G},\boldsymbol{\eta}\right)
\nonumber\\[0.3em]
=& p^T_{ijks}
\nonumber\\
&\, \, + \Bigl(1-p^T_{ijks}\Bigr) \nonumber \\
& \quad \times \underset{\left(\iota \kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }{\Pi}\left(1-\left(1-p^L_{\iota \kappa s}\right)\underset{\mu\in\varOmega_{s}\setminus\{i,j,k\}}{\Pi}\left(1-p^T_{\iota \kappa \mu s}\right)\right)
\end{align}
\subsection{Probability of forming links and triangles, conditional on instruments and first-stage errors}
Next we derive expressions for the probability of seeing links and triangles conditional on the vector of instruments and first-stage errors.We start by inspecting the conditional probabilities $p^L_{\iota \kappa s}$ and $p^T_{\iota \kappa \mu s}$. Note that these are random variables due to their dependence on $\boldsymbol{G},\boldsymbol{\eta}$.
Using the definition of $\xi_{ijs}$ and $\xi_{ijks}$ as well as (ref) and (ref) we can write:
\begin{align*}
p^L_{ijs}&=\beta_{0}+\beta_{1}G_{ijs}+h_{L}(u_{ijs})+\xi_{ijs} \\
p^T_{ijks}&=\theta_{0}+\theta_{1}\sum_{\left(\iota\kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }G_{\iota\kappa s}+h_{T}(u_{ijs},u_{iks,}u_{jks})+\xi_{ijks}
\end{align*}
From the first-stage, $G_{ijs}$ is a deterministic function of $\boldsymbol{Z}$ and $\boldsymbol{u}$. Combining this with the (ref), it follows that $\left\{ p^L_{ijs}\right\} _{(i,j)\in \Omega_s^L} \cup \left\{ p^T_{ijks}\right\} _{(i,j,k)\in \Omega_s^T}$ is mutually independent conditional on $\boldsymbol{Z},\boldsymbol{u}$. Defining $\tilde{p}^L_{ijs}=E[p^L_{ijs}|\boldsymbol{Z},\boldsymbol{u}]$ and $\tilde{p}^T_{ijks}=E[p^T_{ijks}|\boldsymbol{Z},\boldsymbol{u}]$ this implies the following for the conditional expectations of observed link and triangle indicators:
\begin{align}
E\left[L_{ijs}|\boldsymbol{Z},\boldsymbol{u}\right] = & E\left[P\left(L_{ijs}=1|\boldsymbol{G},\boldsymbol{\eta}\right)|\boldsymbol{Z},\boldsymbol{u}\right] \nonumber \\
= &\tilde{p}^L_{ijs}+\left(1-\tilde{p}^L_{ijs}\right)\left(1-\underset{k\in\varOmega_{s}\setminus\left\{ i,j\right\} }{\Pi}\left(1-\tilde{p}^T_{ijks}\right)\right)&
\end{align}
\begin{align}
E\left[T_{ijks} \mid \boldsymbol{Z},\boldsymbol{u}\right] =& E\left[ P\!\left(T_{ijks}=1 \mid \boldsymbol{G},\boldsymbol{\eta}\right) \mid \boldsymbol{Z},\boldsymbol{u} \right]
\nonumber \\
=& \tilde{p}^T_{ijks}
\nonumber\\
&\, \, + \Bigl(1-\tilde{p}^T_{ijks}\Bigr) \nonumber \\
& \quad \times \underset{\left(\iota \kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }{\Pi}\left(1-\left(1-\tilde{p}^L_{\iota \kappa s}\right)\underset{\mu\in\varOmega_{s}\setminus\{i,j,k\}}{\Pi}\left(1-\tilde{p}^T_{\iota \kappa \mu s}\right)\right)
\end{align}
\subsection{Parametric restriction on the conditional expectation of the unobservables}
To set up our control function estimation, we next restrict the conditional expectation of the unobservables to be parametric functions, with parameter vectors $\boldsymbol{\rho}^L$ and $\boldsymbol{\rho}^T$ :
$$h_{L}(u_{ijs})=\tilde{h}_{L}(u_{ijs};\boldsymbol{\rho}^L)-E\left[\tilde{h}_{L}(u_{ijs};\boldsymbol{\rho}^L)\right]$$
$$h_{T}(u_{ijs},u_{iks},u_{jks})=\tilde{h}_{T}(u_{ijs},u_{iks},u_{jks};\boldsymbol{\rho}^T)-E\left[\tilde{h}_{T}(u_{ijs},u_{iks},u_{jks};\boldsymbol{\rho}^T)\right]$$
Note that subtracting off expectated value in these definitions is simply a convenient way to impose a mean of zero.
With these parametric restrictions we can write:
\begin{align}
\widetilde{p}^L_{ijs}&=\beta_{0}+\beta_{1}G_{ijs}+h_{L}(u_{ijs}) \nonumber \\
&=\psi_L +\beta_1 G_{ijs} +\tilde{h}_{L}(u_{ijs};\boldsymbol{\rho}^L) \nonumber\\
&=f_L(G_{ijs},u_{ijs};\psi_L,\beta_1,\boldsymbol{\rho}^L)
\end{align}
the second to last line defines the parameter $\psi_L=\beta_{0}-E\left[\tilde{h}_{L}(u_{ijs};\boldsymbol{\rho}^L)\right]$ while the last line defines the function $f_L$. Completely analogously, we can write:
\begin{align}
\widetilde{p}^T_{ijks} &=\theta_{0}+\theta_{1}\sum_{\left(\iota\kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }G_{\iota\kappa s}+h_{T}(u_{ijs},u_{iks},u_{jks})\nonumber\\
&=\psi_T+\theta_{1}\sum_{\left(\iota\kappa\right)\in\left\{ (i,j),(i,k),(j,k)\right\} }G_{\iota\kappa s}+\tilde{h}_{T}(u_{ijs},u_{iks},u_{jks};\boldsymbol{\rho}^T) \nonumber\\
&=f_T(G_{ijs},G_{iks},G_{jks},u_{ijs},u_{iks},u_{jks};\psi_T,\theta_1,\boldsymbol{\rho}^T)
\end{align}
\subsection{Control function estimation}
Combining (ref), (ref), (ref) and (ref) we can write the conditional expectations $E\left[L_{ijs} \mid \boldsymbol{Z}, \boldsymbol{u} \right]$ and $E\left[T_{ijks} \mid \boldsymbol{Z}, \boldsymbol{u}\right]$ as functions of first-stage error terms, group membership and parameters:
\begin{align*}
E\left[L_{ijs} \mid \boldsymbol{Z}, \boldsymbol{u} \right] & =g_L \left(i,j,s,\boldsymbol{G},\boldsymbol{u};\psi_L,\psi_T,\beta_1,\theta_1,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T)\right) \\
E\left[T_{ijks} \mid \boldsymbol{Z}, \boldsymbol{u}\right] & = g_T \left(i,j,k,s,\boldsymbol{G},\boldsymbol{u};\psi_L,\psi_T,\beta_1,\theta_1,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T \right)
\end{align*}
where
\begin{align}
g_L & \left(i,j,s,\boldsymbol{G},\boldsymbol{u};\psi_L,\psi_T,\beta_1,\theta_1,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T)\right) \nonumber \\ &= f_L(G_{ijs},u_{ijs};\psi_L,\beta_1,\boldsymbol{\rho}^L)\nonumber \\
& \qquad\qquad+\left(1-f_L(G_{ijs},u_{ijs};\psi_L,\beta_1,\boldsymbol{\rho}^L)\right) \nonumber \\ & \qquad \qquad \qquad \times \left(1-\underset{k\in\varOmega_{s}\setminus\left\{ i,j\right\} }{\Pi} \left( 1- f_T(G_{ijs},G_{iks},G_{jks},u_{ijs},u_{iks},u_{jks};\psi_T,\theta_1,\boldsymbol{\rho}^T)\right) \right)
\end{align}
and
\begin{align}
g_T & \left(i,j,k,s,\boldsymbol{G},\boldsymbol{u};\psi_L,\psi_T,\beta_1,\theta_1,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T \right) \nonumber
\\
&= f_T\bigl(
G_{ijs}, G_{iks}, G_{jks},
u_{ijs}, u_{iks}, u_{jks};
\psi_T, \theta_1, \boldsymbol{\rho}^T
\bigr) \nonumber\\
&\quad + \Bigl(1 - f_T\bigl(
G_{ijs}, G_{iks}, G_{jks},
u_{ijs}, u_{iks}, u_{jks};
\psi_T, \theta_1, \boldsymbol{\rho}^T
\bigr)\Bigr) \nonumber\\
&\qquad \times
\Pi_{(\iota,\kappa)\in\{(i,j),(i,k),(j,k)\}}
\Biggl[
1 - \Bigl(1 - f_L\bigl(
G_{\iota\kappa s}, u_{\iota\kappa s};
\psi_L, \beta_1, \boldsymbol{\rho}^L
\bigr)\Bigr) \nonumber\\
&\qquad\qquad \times
\Pi_{\mu\in\varOmega_s\setminus\{i,j,k\}}
\Bigl(
1 - f_T\bigl(
G_{\iota\kappa s}, G_{\iota\mu s}, G_{\kappa\mu s},
u_{\iota\kappa s}, u_{\iota\mu s}, u_{\kappa\mu s};
\psi_T, \theta_1, \boldsymbol{\rho}^T
\bigr)
\Bigr)
\Biggr]
\end{align}
This leads to the two-step control function estimator that we use which entails the following two steps:
\begin{enumerate}
• Estimate the first stage regression, (ref), using data on all pairs, to obtain the residuals $\widehat{\boldsymbol{u}}$.
• Use the residuals as estimates of the true first stage errors and estimate the parameters of interest $\beta_1,\theta_1$ as well as $\psi_L,\psi_T,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T$ using (weighted) nonlinear least squares on the sample of all pairs and all triads:
$$(\hat{\beta}_1,\hat{\theta}_1,\hat{\psi}_L,\hat{\psi}_T,\widehat{\boldsymbol{\rho}}^L,\widehat{\boldsymbol{\rho}}^T)=\textrm{argmin}_{\beta_1,\theta_1,\psi_L,\psi_T,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T}\quad Q(\beta_1,\theta_1,\psi_L,\psi_T,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T)$$
where
\begin{align}
Q(&\beta_1,\theta_1,\psi_L,\psi_T,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T) \\
=&\sum_s\sum_{ij\in\varOmega_s^L} \left(L_{ijs}-g_L \left(i,j,s,\boldsymbol{G},\widehat{\boldsymbol{u}};\psi_L,\psi_T,\beta_1,\theta_1,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T\right) \right)^2 \nonumber \\
& \qquad +\omega_T \sum_s \sum_{ijk\in\varOmega_s^T} \left(T_{ijks}- g_T \left(i,j,k,s,\boldsymbol{G},\widehat{\boldsymbol{u}};\psi_L,\psi_T,\beta_1,\theta_1,\boldsymbol{\rho}^L,\boldsymbol{\rho}^T \right) \right)^2 \nonumber
\end{align}
\end{enumerate}
In the above, $\omega_T$ is a deterministic constant that governs how much weight is given to triads in estimation. For our baseline estimates we simply set $\omega_T=1$ (another obvious choice would be to set it to the ratio of pairs to triads in the data so as to give equal weight, $\omega_T=\frac{\sum_s |\varOmega_s^L|}{\sum_s |\varOmega_s^T|}$). To obtain standard errors and conduct inference, we apply a bootstrap procedure that resamples study programs and repeats both steps above on the the bootstrap samples.
In implementing the estimator, we of course need to settle on a particular functional form for $\tilde{h}_L$ and $\tilde{h}_T$. A natural baseline choice would be second order polynomials that are symmetric in the arguments. This corresponds to setting $\boldsymbol{\rho}^L=(\rho_1^L,\rho_2^L)$,$\boldsymbol{\rho}^T=(\rho_1^T,\rho_2^T,\rho_3^T)$ and setting
\begin{align*}
\tilde h_L(u;\boldsymbol{\rho}^L)&=\rho_1^L u+\rho_2^L u^2
\\
\tilde h_T(u_1,u_2,u_3;\boldsymbol{\rho}^T)
&=
\rho_1^T(u_1+u_2+u_3)
+\rho_2^T(u_1^2+u_2^2+u_3^2)
+\rho_3^T(u_1u_2+u_1u_3+u_2u_3)
\end{align*}
Other choices could be simply (symmetric) linear functions
\begin{align*}
\tilde h_L(u;\boldsymbol{\rho}^L)&=\rho_1^L u
\\
\tilde h_T(u_1,u_2,u_3;\boldsymbol{\rho}^T)&=\rho_1^T(u_1+u_2+u_3),
\end{align*}
with $\boldsymbol{\rho}^L=(\rho_1^L)$ and $\boldsymbol{\rho}^T=(\rho_1^T)$, or a higher-order approach based on monomials such as
\begin{align*}
\tilde h_L(u;\rho_L)=&\rho_1^L u+\rho_2^L u^2+\rho_3^L u^3
\\
\tilde h_T(u_1,u_2,u_3;\rho^T)
=
&\rho_1^T(u_1+u_2+u_3)
+\rho_2^T(u_1^2+u_2^2+u_3^2)
+\rho_3^T(u_1u_2+u_1u_3+u_2u_3)
\\
&\quad +\rho_4^T(u_1^3+u_2^3+u_3^3)
\\
&\qquad+\rho_5^T(u_1^2u_2+u_1^2u_3+u_2^2u_1+u_2^2u_3+u_3^2u_1+u_3^2u_2)
\\
& \qquad \quad +\rho_6^T u_1u_2u_3
\end{align*}
with $\boldsymbol{\rho}^L=(\rho_1^L,\rho_2^L,\rho_{3}^L)$ and $\boldsymbol{\rho}^T=(\rho_1^T,\rho_2^T,\rho_3^T,\rho_4^T,\rho_5^T,\rho_6^T)$.
Note that after implementing the estimator above, we can also recover estimates of the constant terms in the original link and triangle formation equations by simply inverting the definitions of $\psi_L$ and $\psi_T$, and plugging in estimates and sample averages for expectations and true parameters:
\begin{align*}
\hat\beta_0
& =
\hat\psi_L
+
\frac{1}{\sum_s |\Omega_s^L|}
\sum_s \sum_{(i,j)\in \Omega_s^L}
\tilde h_L(\hat u_{ijs};\hat{\boldsymbol{\rho}}^L),
\\
\hat\theta_0
& =
\hat\psi_T
+
\frac{1}{\sum_s |\Omega_s^T|}
\sum_s \sum_{(i,j,k)\in \Omega_s^T}
\tilde h_T(\hat u_{ijs},\hat u_{iks},\hat u_{jks};\hat{\boldsymbol{\rho}}^T)
\end{align*}
\putbib[bibliography]
\part*{Supplementary Material}
\addcontentsline{toc}{part}{Supplementary Material}