EconBase
← Back to paper

The direct and spillover effects of a nationwide socio-emotional learning program for disruptive students

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

109,142 characters · 24 sections · 110 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The direct and spillover effects of a nationwide socio-emotional learning program for disruptive students.

abstractSocial and emotional learning (SEL) programs teach disruptive students to improve their classroom behavior. Small-scale programs in high-income countries have been shown to improve treated students’ behavior and academic outcomes. Using a randomized experiment, we show that a nationwide SEL program in Chile has no effect on eligible students. We find evidence that very disruptive students may hamper the program's effectiveness. ADHD, a disorder correlated with disruptiveness, is much more prevalent in Chile than in high-income countries, so very disruptive students may be more present in Chile than in the contexts where SEL programs have been shown to work.

Keywords: disruptive students, spillover effects, peer effects, social and emotional learning.

\vskip\medskipamount JEL Codes: I21, I24, I28, D62.

\let\markeverypar\everypar \newtoks\everypar \everypar\markeverypar \markeverypar{\the\everypar\looseness=-2\relax} \thispagestyle{empty} \parskip=3.pt \baselineskip=20pt \linespread{1}

Introduction

\setcounter{page}{1}

lazear2001 has proposed that classroom learning is a public good suffering from congestion effects, which are negative externalities created when one student is disruptive and impedes the learning of her classmates. In the US, those externalities are important: carrell2010externalities and carrell2018long find that being exposed to one peer experiencing domestic violence at home, a good proxy for a disruptive peer, reduces classmates' test scores by 0.07 standard deviation ($\sigma$), and reduces their earnings at age 26 by 3 to 4 percent. figlio2007 also finds that being exposed to disruptive peers reduces classmates test scores. betts1999 find that US middle and high schools teachers devote 6.1% of instruction time to discipline, and that this fraction is higher in disadvantaged schools. Therefore, programs effective at reducing troubled students' disruptiveness may generate large positive spillover on their classmates, on top of their direct effects.

\vskip\medskipamount Epidemiological studies show that the prevalence of ADHD, a disorder correlated with conduct problems, is higher in some low- and middle-income countries than in high-income countries. Then, addressing students' conduct problems may be an even more pressing issue in those countries. In Chile, the country where the intervention we study takes place, 15.5% of primary school children have ADHD (see de2013epidemiology). Primary school children also have be found to have high ADHD rates in Colombia (16.9%, see cornejo2005prevalencia), or in Iran (17.3%, see safavi2016prevalence). On the other hand, the ADHD prevalence rate among primary school children is estimated at 6.8% in the US (see visser2014trends), between 3.5 and 5.6% in France (see lecendreux2011prevalence), and 3% in Italy (see bianchini2013prevalence).

\vskip\medskipamount School-based mental health programs are often used to reduce students' disruptiveness. Some programs are universal, meaning that they are delivered in classroom settings to all the students in the class. Other programs are selected, meaning that they are provided to students identified by teachers as having conduct problems, during the school day and outside the classroom. Many school-based mental health programs are social and emotional learning (SEL) programs (see wilson2007school), that teach children to recognize and manage their emotions, and to handle interpersonal situations effectively, using cognitive and behavioral therapy (CBT). A vast literature has found SEL programs to be successful. In a meta-analysis of 80 selected interventions, payton2008 find that they reduce conduct problems by $0.47\sigma$, and respectively improve mental health and academic performance by $0.50$ and $0.43\sigma$. Meta-analyses of universal SEL interventions find smaller but still large effects on those dimensions, around $-0.25\sigma$ for conduct problems, and $+0.55\sigma$ and $+0.30\sigma$ for mental health and test scores (durlak2011impact, sklad2012effectiveness, wigelsworth2016impact, taylor2017promoting, and corcoran2018effective).

\vskip\medskipamount However, there are at least two gaps in the literature. First, it has mostly considered small-scale demonstration programs mounted by researchers in a handful of schools. The effect of SEL interventions may differ when implemented at scale (see davis2017 or weisz2014). Second, it has mostly focused on interventions conducted in high-income countries, while epidemiological studies suggest that addressing students' conduct problems may be more pressing in some middle- and low-income countries. A recent meta-analysis of psycho-social interventions for disruptive students in low- and middle-income countries (see burkey2018psychosocial) includes only two SEL interventions, one in Jamaica and the other in Romania. Both are universal interventions, implemented at a very small scale. Both find positive effects on students' conduct problems (see baker2009pilot and cstefan2013).

\vskip\medskipamount This paper contributes to addressing those two gaps: it is the first to measure the effects of a nationwide SEL program, and the program we study takes place in a middle-income country with a high ADHD prevalence rate. Specifically, we study the effects of “Skills for Life” (SFL), a selected SEL program for disruptive second graders in Chile. Since its creation in 1998, SFL has screened and treated around 1,000,000 children, making it the fifth largest school-based mental health program in the world (see murphy2017). To identify eligible students, SFL teams use a psychometric scale measuring students' disruptiveness, and students above some cut-off are eligible. Eligible students then follow 10 two-hours SEL sessions with a psychologist and a social worker. SFL is a costly program: we estimate that its cost per student is equivalent to 15% of the expenditure per primary school student in Chile.

\vskip\medskipamount We randomly assigned 172 classes to either receive SFL in the first or second semester of the 2015 school year, and we measured outcomes at the start of the second semester, after the treatment group had received the treatment but before the control group received it. By comparing eligible students in the treatment and control groups, we can estimate the direct effects of the program, and by comparing ineligible students in the two groups we can estimate its spillover effects.

\vskip\medskipamount We find that SFL does not have effects on eligible students' disruptiveness, mental health, and academic achievement. The effects we can rule out are fairly small, and much smaller than those found in payton2008 and in all the other SEL meta-analyses we are aware of. For instance, we can rule out at the 5% level that the program increases students’ Spanish scores by more than 0.09$\sigma$, or that it reduces teachers’ assessment of students’ disruptiveness by more than 0.10$\sigma$. Not surprisingly, as we do not find that SFL impacts eligible students, we also do not find spillover effects on ineligible ones. Finally, we even find that the program has a strong negative effect on teachers' and enumerators' ratings of the overall disruptiveness of treated classes.

\vskip\medskipamount To account for the discrepancy between our results and the literature, we compared SFL to the selected SEL interventions reviewed in payton2008. Three conclusions emerge. First, SFL's intensity (number of sessions, duration...) is comparable to that of the meta-analysis's interventions, so it is not the case that SFL is not intensive enough to produce an effect. Second, as ADHD is much more prevalent in Chile than in high-income countries, SFL may be faced with a harder-to-treat population than the interventions reviewed in payton2008. We indeed find evidence that SFL's effectiveness is hampered by the presence of very disruptive students. In classes with at least one very disruptive eligible student, defined as students above eligible students' 90th percentile of baseline disruptiveness,\footnote{This definition ensures that about 50% of classes have at least one very disruptive student.} the program increases the disruptiveness of other eligible students, of ineligible students, and it worsens teachers' and enumerators' ratings of the overall class disruptiveness. The program also strongly increases the friendship ties between very disruptive and other eligible students. The former may then have a negative influence on the latter, which would explain the negative effects we observe. Third, SFL and the meta-analysis's interventions strikingly differ in terms of scale and delivery. The interventions in the meta-analysis are demonstration programs mounted by researchers, that typically treat a few dozens children in a handful of schools. Half are delivered by the researchers, while the other half are delivered by psychologists or teachers under researchers' close supervision: typically, researchers review their delivery of the intervention every week. On the other hand, SFL is a large-scale governmental program, delivered by psychologists without any researcher involvement. The governmental agency in charge of the program loosely monitors the program implementers, and very rarely audits their workshops. Without sufficient monitoring, teams may not implement the program with high-enough fidelity, which could also explain why SFL does not produce an effect.

\vskip\medskipamount The remainder of the paper is organized as follows. In Section (ref), we present the SFL program. In Section (ref), we present the randomization, the data we use, and the population under study. In Section (ref), we present compliance with randomization, the balancing checks, and attrition. In Section (ref), we present the main results. In Section (ref), we interpret the results and present some exploratory analysis.

The SFL program

SEL is the process through which children acquire the skills to recognize and manage their emotions, set and achieve positive goals, and handle interpersonal situations effectively. SEL programs try to enhance children's self-awareness (accurately assessing one's feelings and maintaining a sense of self-confidence), self-management (regulating one's emotions and controlling impulses), and social awareness (being able to take the perspective of others, preventing, managing, and resolving interpersonal conflict). Selected programs are provided to specific students identified as having conduct problems, during the school day and outside of their classroom. Meta-analysis have shown that selected SEL programs improve SEL skills, reduce conduct problems, and can improve academic achievement (see payton2008).

\vskip\medskipamount SFL is a Chilean school-based selected SEL program for second graders suffering from conduct disorders. It is managed by JUNAEB (Junta Nacional de Auxilio Escolar y Becas), the division of the Chilean Department of Education in charge of most of the non-teaching programs implemented in Chilean schools. The program started as a pilot in 1998. Over the next 3 years, JUNAEB collaborated with psychologists from the University of Chile to review the screening measures and programs available at that time, and design an SEL program adapted to Chile, where the prevalence of ADHD is particularly high among children (see de2013epidemiology). The program became a nationwide policy in 2001, and it is currently implemented in 1,637 publicly-funded elementary schools in Chile (see guzman2015). These schools account for 20% of all elementary schools in Chile, and they are the most disadvantaged. Since 1998, the SFL program has screened and treated around 1,000,000 children, making it the fifth largest school-based mental health program in the world (see murphy2017).

\vskip\medskipamount To identify eligible students, SFL uses a psychometric scale, the Teacher Observation of Classroom Adaptation (TOCA, see kellam1977, and werthamer1990), adapted to the Chilean context by george1994adaptacion. In the end of each academic year, first-grade teachers answer the TOCA questionnaire for each of their student. Based on this questionnaire, students receive scores on the following six scales: authority acceptance (AA), attention and focus (AF), activity levels (AL), social contact (SC), motivation for schooling (MS), and emotional maturity (EM). The TOCA questionnaire concludes with two summary questions, where teachers have to give ratings of the overall disruptiveness and academic ability of each of their student.

\vskip\medskipamount Then, the three following groups of students are eligible for the program:

itemize• Students above the 75th percentile of the AA scale, above the 85th percentile of the AF and AL scales, and below the 25th percentile of the MS scale; • Students below the 25th percentile of the SC scale, and either above the 75th percentile of the AA scale or above the 85th percentile of the AL scale; • Students below the 25th percentile of the SC, MS, and EM scales, and below the 50th percentile of either the AA or AL scale.

The percentiles are gender specific, to ensure that not only males are eligible, and were computed using a representative sample of the 2nd grade population in Chile (see george1994adaptacion,de2005prediction). Students in the third eligibility group are not disruptive, but they only account for 7% of eligible students, while the first two groups respectively account for 40% and 53% of eligible students. Depending on the year, eligible students account for 15 to 20% of first-grade students whose teachers answer the TOCA questionnaire.

\vskip\medskipamount In second grade, SFL asks eligible students' parents the authorization to enroll their child in the program. If their parents accept, eligible students are enrolled in a workshop implemented by a team of two SFL employees. A survey conducted in 2015 (see rojas2018) shows that half of SFL employees are psychologists. In Chile, this title can be obtained after a college degree with a psychology major (see guzman2015). The other half of employees are social workers and former teachers, titles that can also be obtained after a college degree. Usually, an SFL team consists of a psychologist and a social worker or teacher. 77% of SFL employees are women, their average age is 31 years old. They have on average 2.6 years of experience into the program, and 36% have less than one year of experience, indicating a high rate of turnover. During their first year, SFL employees receive three eight-hours-long days of training (see rojas2017efectos). They also attend “good practices” meetings every six months, in which they share with other teams what works in their workshops. As the Chilean public school system is administrated at the municipal level, SFL teams are also organized at this administrative level.

\vskip\medskipamount SFL workshops consist in 10 two-hours group sessions, taking place weekly, during the class day, over the course of one semester. During sessions, enrolled students leave the classroom, while their classmates stay there and continue with their normal schedule. The time of the group sessions is set in coordination with teachers, to avoid that enrolled students lose key instruction time. During the workshops, teachers teach subjects deemed less crucial than Spanish or mathematics, like religion (a mandatory subject in Chile) or music, to the ineligible students that stay with them in the classroom (see rojas2018). The workshop takes place over two school periods, and eligible students come back to their classroom before the break.

\vskip\medskipamount Sessions are divided into five parts. The goals of the first part are to welcome children and build a group identity, for instance by having children choose a group name. The goals of the second part are to improve children's self-esteem, and their respect of others. Then, during the third part, the psychologists help students put words on their and others' emotions, and help them share their emotions with others. Then, the fourth part is dedicated to self-control techniques, and to strategies to find non-violent solutions to conflicts. Finally, the last part is dedicated to a review of what has been learnt during the workshop. Sessions are activity based, involve games and role play, and make use of CBT techniques. If they behave well during a session, students sometimes receive rewards like cakes or candies. SFL employees are provided with a 114-pages-long manual describing the goal and the content of each session, and suggesting games and activities. But they are also encouraged to tailor the content of their sessions to the specific needs of the students enrolled.

\vskip\medskipamount As per the SFL guidelines, six to 12 students should participate in a workshop. If there are less than six eligible students in a school, no workshop takes place, and if a school has more than 12 eligible students, two workshops take place in that school. In the next section, we explain how we exploit these features in our randomization. Finally, the parents of enrolled children are invited to three training sessions, whose goal is to encourage them to reproduce the workshop's activities at home.

\vskip\medskipamount We estimate that SFL costs 200 USD per treated student. We also estimate that the government spends 1,316 USD on instruction per student and per year in the schools in our sample.\footnote{The government funds public schools by giving them a voucher per student, whose amount depends on the student's attendance (\url{https://www.oecd-ilibrary.org/docserver/9789264287112-6-es.pdf?expires=1586606397&id=id&accname=guest&checksum=17AF0B3C9CF0863F8300FAA082FE969D}). For public primary schools, the school voucher is worth 754 USD for an attendance of 84%, the average attendance observed in our sample. Then, the government gives schools an extra voucher worth 721 USD for every very disadvantaged student (\url{https://ate.mineduc.cl/usuarios/admin3/doc/2015020312570909985.Manual_Apoyo_a_la_Gestion.pdf}), and 78% of students are very disadvantaged in the schools we study, thus leading to our 754+0.78$\times$721=1,316 USD estimate.} Therefore, the program's cost represents a sizeable 15% increase of the expenditure per student. JUNAEB does not have an estimate of the total cost of the program, here is how we estimated it. The 2014 budget of one of the municipal teams in our sample shows that its program implementers earned on average 7.42 USD per hour in 2014. Then, based on interviews with two implementers, we estimated that it takes 149 hours of work to implement an SFL workshop. This includes the 52 hours that the two workshop implementers spend delivering 13 two-hours sessions to students and their parents, but also the time that they spend: preparing the sessions and buying the material they need; going to and returning from the school for each session; preparing the reporting documents JUNAEB asks them to send for each workshop; meeting with the school principal and 2nd grade teachers prior to the start of the workshop, to agree on the schedule and location of the workshop; and interview 1st grade teachers to fill the TOCA questionnaire for each of their students the year before the workshop. Then, the team's budget shows that digitizing the 2014 TOCAs of all the first grade students in the town costed 860 USD. Divided by the 10 workshops conducted that year, that leads to a cost of 86 USD per workshop. Implementers also received transportation vouchers worth 63 USD per workshop. Finally, the cost of the material needed for the workshop activities is estimated at 188 USD per workshop, based on a detailed list of all the items bought for a workshop provided by the implementers we interviewed. Overall, we estimate the total cost of a workshop at 7.42$\times$149+86+63+188=1,443 USD. The team whose budget we used had 7.2 students per workshop in 2014, which finally yields our estimated cost of 200 USD per treated student. This estimate relies on one team's budget. Costs may vary between teams, but we do not have reasons to suspect that the program's average cost is orders of magnitude away from our estimate.

\vskip\medskipamount Previous research has found that from first to third grade, the disruptiveness of students that attend seven to 10 SFL sessions in second grade decreases more than that of students attending six sessions or less (see e.g. guzman2015). However, SFL attendance is driven by students' school attendance, and students who attend school less may do so because they experience negative shocks, which could explain why their disruptiveness decreases less. To avoid that type of endogeneity bias, our paper relies on an experimental control group to measure the effect of SFL.

Randomization, data, and study population

Sample selection and randomization

Our sample consists of 172 classes. All municipal teams conducting the SFL program in the Santiago and Valparaiso regions, the two most populated regions in Chile, were invited to join the study. 32 out of 39 accepted our invitation. In March 2015, these teams visited the schools covered by the program in their municipalities, and collected data on the number of students eligible for the program enrolled in each second grade class. 172 classes with four or more eligible students and in schools with six or more eligible students were included in the study. The second criterion ensured that group sessions would indeed take place in the school, while the first criterion ensured that there were enough treated students per class to potentially generate spillover effects. About 450 classes participate in a SFL workshop each year in the Santiago and Valparaiso regions, so our sample covers about 40% of the classes covered by the program in those regions.

\vskip\medskipamount Randomization took place both within schools and within municipalities. There were 29 schools with two classes included in our sample and where it was possible to form two groups of six students or more without grouping students of the two classes together. In such instances, we conducted a lottery within the school, to assign one of the two classes to receive the treatment in the first semester of 2015, and the second class to receive it in the second semester. The remaining 114 schools each only had one class included in our sample, so randomization took place within municipalities. In this latter group of schools, there is no risk that the control group students may have been contaminated by the treatment, while this may have happened in the former group of schools. Later in the paper, we reestimate the treatment effect in the second group of schools and find very similar effects to those we find in the full sample, so control-group contamination does not seem to drive our results.

\vskip\medskipamount Overall, we conducted 56 lotteries (29 within schools, and 27 within municipalities) and we assigned 89 classes to receive the treatment in the first semester, from April to June 2015, and 83 to receive it in the second semester, from September to December 2015.

Data

In our analysis, we use data produced by JUNAEB. First, we use the six first-grade TOCA scores that determine students' eligibility to SFL, as well as the teachers' ratings of students' disruptiveness and academic ability in the TOCA questionnaire. Then, we also use another psychometric scale collected by JUNAEB and measuring students' disruptiveness, the pediatric symptom checklist (PSC, see jellinek1988), which is filled by students' parents. We also use JUNAEB's data on treatment implementation. Specifically, for each class in our sample we know how many SFL group sessions were conducted in the first semester of 2015. For each student, we know how many sessions she attended, and how many sessions her parents attended. Finally, JUNAEB also provided us data on students' socio-economic background, as well as their monthly school attendance from March 2015 to June 2015.

We also use baseline data collected in March 2015, before the treatment started in the treatment group classes, and endline data collected in August 2015, after the treatment ended in the treatment group classes and before it started in the control group classes. Both at baseline and endline, two enumerators visited each of the 172 classes included in the experiment during a half day. Enumerators were undergraduate students, mostly psychology and education majors. Every person who applied to become an enumerator first had to attend a half-day training, during which he/she was taught how to administer our questionnaires. Candidates also had to take a test at the end of the training, and only those who scored above some threshold became enumerators.

\vskip\medskipamount Our questionnaires slightly changed from baseline to endline. Below, we describe our endline questionnaires, and we explain the difference between our baseline and endline questionnaires when needed later in the paper.

\vskip\medskipamount The enumerators first administered a non-cognitive questionnaire to the students. That questionnaire aimed at measuring:

itemize• Students' happiness in school, using a question from the student SIMCE questionnaire.\footnote{The SIMCE (Sistema de Medici{\'o}n de la Calidad de la Educaci{\'o}n) questionnaires are the nationwide standardized cognitive and non cognitive questionnaires administered to students and teachers in Chile.} • Students' self-control, using items of the child self-control psychometric scale (see rorhbeck1991child) that we translated into Spanish. • Students' self-esteem, using items of the self-perception for children psychometric scale (see harter1985manual) translated and validated into Spanish (see molina2011adaptacion).

\vskip\medskipamount Second, the enumerators administered a Spanish and mathematics test to the students. Third, the enumerators interviewed individually each student and asked her to name up to three students that she likes to play with during breaks, hereafter referred to as the student's friends. Fourth, the enumerators observed a one-hour lecture. During that observation, they observed the behaviour of each student during five seconds, and assessed whether the student was studying, not studying, or being disruptive. They repeated that process five times, and then rated the overall disruptiveness of each student by answering the summary question from the TOCA questionnaire. During that one-hour lecture, the enumerators also recorded the decibel levels in the class using a smartphone app, and wrote down the time at which the lecture was supposed to start and the time when it effectively started. Fifth, the enumerators filled a short questionnaire aimed at assessing the overall disruptiveness in the class, using questions taken from the PISA (Program for International Student Assessment) questionnaire, asking them their agreement with statements such as: “There is noise and disorder in this class,” or “The teacher has to wait for a long time before students calm down and he/she can start teaching”.

\vskip\medskipamount The enumerators also administered a questionnaire to the teachers. That questionnaire aimed at collecting: teachers' socio-demographic characteristics; teachers' ratings of the overall disruptiveness of the class, using similar questions as those asked to enumerators; teachers' rating of the prevalence of bullying in the class; teachers' motivation, taste for their job, and mental health levels. The questionnaire was for the most part composed of questions from the SIMCE teacher questionnaire. Teachers also rated the overall disruptiveness of each of their student by answering the summary question from the TOCA questionnaire.

\vskip\medskipamount Finally, in July 2019 we also conducted qualitative interviews to shed light on the mechanisms underlying our results. We interviewed three of the SFL municipal teams that had participated in our experiment, and that account for 12% of our sample.

\vskip\medskipamount The list of the outcome variables we consider in the paper was pre-specified in a pre-analysis plan (PAP) available at \url{https://www.socialscienceregistry.org/trials/1080}. That plan was time-stamped on 04/28/2017, before JUNAEB sent us students' first grade TOCA scores, as a letter from JUNAEB officials also available on the social science registry website testifies. Students' first-grade TOCA scores are necessary to distinguish eligible and ineligible students in our data, a distinction that underlies most of our analysis. Even though endline took place almost two years before we submitted our PAP, we had not started to analyze our data before. Indeed, we had not finished cleaning the data before submitting our PAP. This research was funded through four small grants, totalling 37,000 GBP. Therefore, we could not afford to buy tablets to collect our data, and instead used paper questionnaires. We could also not afford to hire a RA in charge of supervising data entry and data cleaning. Instead, we supervised the RAs in charge of data entry ourselves, and we also took care of the data cleaning ourselves. This process ended after 04/28/2017.

\vskip\medskipamount The analysis presented in Sections (ref) and (ref) follows our pre-analysis plan, except for a few exceptions described below. On the other hand, the analysis presented in Section (ref) was not pre-specified in our PAP. The student-level outcome measures listed in our PAP are:

itemize• the student's happiness in school, self-control, self-esteem, Spanish, and mathematics scores, • the percentage of school days missed by the student from April to June 2015, • the rating of the student's disruptiveness by her teacher, • the average rating of the student's disruptiveness across the two enumerators, • the percentage of the student's classmates that nominate her as one of their friends, • an indicator for whether the student is not nominated as a friend by any other student, • the average disruptiveness at baseline of the student's endline friends, • the average baseline Spanish and mathematics scores of the student's endline friends.

The class-level outcome measures listed in our PAP are:

itemize• the teacher's rating of the class's disruptiveness, constructed using teachers' answers to the PISA questions measuring the disruptiveness in the class, • the teacher's rating of the prevalence of bullying in the class, • the average rating of the class's disruptiveness across the two enumerators, constructed using enumerators' answers to the PISA questions measuring the disruptiveness in the class, • the number of minutes between the moment the class was supposed to start and the moment it effectively started according to the enumerators, • the average decibel levels during the class across the two enumerators' recordings.

We standardize the school happiness, self-control, self-esteem, disruptiveness and test score measures to have a mean of 0 and a $\sigma$ of 1 in the sample.

Assessing data quality

Some of the dimensions we are trying to measure are hard to observe. To get a sense of the reliability of our measures, Table (ref) shows their baseline-endline correlation in the control group. Students' Spanish and mathematics test scores have high positive baseline-endline correlations, above 0.5. Those correlations are still far from one, probably because students in our study are young and their cognitive ability is not fixed yet. Our measure of students' popularity has a baseline-endline correlation of 0.32. Our school happiness, self-esteem, and self-control measures respectively have baseline-endline correlations of 0.22, 0.13, and 0.14.

\vskip\medskipamount Turning to disruptiveness measures, the rating of students' disruptiveness by teachers has a baseline-endline correlation of 0.42, which is almost as high as the baseline-endline correlation of test scores. This is all the more remarkable as we use first grade teachers' answer to the TOCA summary question as our baseline measure,\footnote{We decided to include the summary TOCA question in our baseline teacher questionnaire after having collected more than half of the baseline data, so that variable is missing for many classes at baseline.} so our baseline and endline measures were not made by the same teacher. This suggests that students' disruptiveness is relatively stable, and that different teachers tend to agree in their ratings. Then, Table (ref) shows that this measure is negatively correlated with students' academic ability: at baseline, its correlation with students' average test score in Spanish and mathematics is equal to -0.28. Finally, the bottom panel of Table (ref) shows that teachers' rating of the disruptiveness of the class also has a high baseline-endline correlation, equal to 0.50.

\vskip\medskipamount In our PAP, we had planned to use the average of the two enumerators' ratings of a student's disruptiveness as our enumerator disruptiveness rating. However, this measure has a baseline-endline correlation close to, and insignificantly different from, zero. This could be due to the fact that endline and baseline observations are made by different enumerators, who may have different standards to assign a given grade on the disruptiveness scale. Therefore, we depart from our PAP, and slightly modify our measure. We start by regressing enumerators' ratings on enumerator fixed effects, in the sample of control group classes. Then, we compute the residuals from that regression both for treatment and control group classes, and we use the average of those residuals, across the two enumerators that have rated a student, as our enumerators' rating. This modified measure is the difference between a student's average rating by the two enumerators and the average of the ratings made by the same enumerators in the control group. Panel A of Table (ref) shows that it has a positive and significant baseline-endline correlation equal to 0.13, and Panel A of Table (ref) shows that it correlates well with teachers' ratings, and reasonably well with students' academic ability. Overall, enumerators' ratings of students' disruptiveness seem noisier than teachers', but they are still meaningful. Then, Panel B of Table (ref) shows that enumerators' ratings of classes' disruptiveness have a relatively high baseline-endline correlation, around 0.25, and Panel B of Table (ref) shows that this measure correlates well with teachers' ratings. Contrary to teachers' ratings, enumerators' ratings are blinded: enumerators do not know if the class they observe has been treated or not.\footnote{Previous literature on SEL interventions has also relied on non-blinded teacher ratings (see payton2008).}

\vskip\medskipamount The decibel measure constructed following our PAP also has a very low baseline-endline correlation, and it does not correlate at all with teachers' and enumerators' ratings of classes' disruptiveness. The app's measurement does not seem very precise: enumerators recording the same lecture sometimes end up with average noise levels differing by more than 10 decibels. This measurement also seems to depend on the make of the phone and on idiosyncratic factors specific to the enumerator's phone. Therefore, we depart again from our PAP, and net out enumerators' fixed effects from decibel measures, exactly as we did for enumerators' disruptiveness ratings. This new measure has a higher baseline-endline correlation than the measure described in our PAP, though Table (ref) shows that this correlation is still not significant. But it also has a much larger correlation with enumerators' ratings of the class disruptiveness, and that correlation is significant as shown in Table (ref).

Study population

The 172 classes included in our sample bear 5,704 students, meaning that classes have an average of 33.2 students. 4,466 students are ineligible to the program (26.0 per class), while 1,238 students are eligible (7.2 per class). Column (1) in Table (ref) below presents the baseline characteristics of ineligible students. 33.8% of them are born to teenage mothers, which is more than twice the corresponding proportion in Chile.\footnote{See \url{http://web.minsal.cl/portal/url/item/c908a2010f2e7dafe040010164010db3.pdf}.} 75.2% of them live in households below the 20th percentile of the social security score. Being below this threshold opens eligibility for 22 social programs and is usually considered as a proxy for poverty. 44.4% of them live in households below the 5th percentile of the social security score. Being below this threshold opens eligibility for 3 more social programs and is usually considered as a proxy for extreme poverty. Overall, the students included in our study live in households disproportionately coming from the bottom of the Chilean income distribution.

\vskip\medskipamount Column (2) in Table (ref) presents the baseline characteristics of eligible students, and Column (3) reports the p-value of tests that the baseline characteristics of eligible and ineligible students are equal. Panel A shows that eligible students are more likely to be males and less likely to live with their father. Their parents are also less educated than that of ineligible students. Panel B shows that eligible students's self-control and self-esteem scores are about 0.2$\sigma$ lower than that of ineligible students. Differences are even more pronounced when one considers students' disruptiveness and academic ability. Eligible students score 1.2$\sigma$ higher than ineligible students on first-grade teachers' disruptiveness ratings, and 0.4$\sigma$ higher on enumerators' baseline ratings. They also score 0.4$\sigma$ lower on the Spanish and mathematics tests. Eligible students are also less popular than ineligible ones: 7.6% of the students in the class nominate them as friends, against 8.8% for ineligible students. The average disruptiveness of their friends is also about 0.2$\sigma$ higher than that of ineligible's friends, thus suggesting some assortative matching along the disruptiveness dimension.

\vskip\medskipamount Finally, Table (ref) shows some characteristics of the teachers in our sample. 96.3% of teachers are females. Their average age is 42.8 years old, they have an average of 16.5 years of experience as a teacher, and 8.6 years of experience in the school where they currently teach.

table[table omitted — 1,940 chars of source]

Compliance, internal validity, and estimation methods

Compliance with randomization and fidelity of treatment assignment

In this section, we show that the SFL teams followed the randomization, and implemented the treatment as per the program's rules: in the treatment group classes, very few ineligible students received the program. To do so, we estimate the effect of being assigned to treatment on actual exposure to treatment during the first semester of 2015. Let $Y_{ijk}$ be a measure of exposure to treatment for student $i$ in class $j$ and lottery $k$. We estimate the following regression:

equation[equation omitted — 76 chars of source]

where the $\gamma_k$s are fixed effects for the 56 lotteries we conducted to assign the treatment, and where $D_{jk}$ is equal to 1 if lottery $k$ assigned class $j$ to the treatment group and to 0 otherwise. $\widehat{\beta}$ estimates a weighted average across lotteries of the within-lottery difference between the average of $Y_{ijk}$ in treatment and control group classes. As our lotteries have few classes, the treatments of classes in the same lottery are strongly negatively correlated. Therefore, we cluster standard errors at the lottery level, following the recommendation of Chaisemartin2019clusterpairRCTs), who show that clustering at the class level could lead to substantial over-rejection of the null hypothesis.

\vskip\medskipamount To estimate the effect of assignment to treatment on class-level measures of exposure, we estimate Regression (ref), except that we use propensity score reweighting instead of lottery fixed effects. With propensity score reweighting, $\beta$ is also identified out of comparisons of treatment and control group classes in the same lottery (see hirano2003). Using propensity score reweighting ensures that the regression does not have too many independent variables with respect to its number of observations (with lottery fixed effects, Regression (ref) would have 57 independent variables and at most 172 observations). In any case, as the share of treated classes is equal to 0.5 in 46 of the 56 lotteries, using lottery fixed effects or propensity score reweighting does not make a large difference.

\vskip\medskipamount Column (1) of Table (ref) below shows the mean value of eight measures of exposure to the treatment in the control group. Column (2) shows estimates of $\beta$ for these eight measures. Column (3) shows estimates of the standard error of $\widehat{\beta}$. Column (4) shows the p-value of a t-test of $\beta=0$. To account for the fact that we consider several measures of exposure to the treatment, Column (5) shows the p-value controlling the False Discovery Rate (FDR) across the eight tests (see benjamini1995controlling). Finally, Column (6) shows the number of observations used in the estimation.

\vskip\medskipamount Panel A of the table shows that SFL sessions were conducted in 8.4% of the control group classes and in 98.1% of the treatment group classes. On average, 0.6 sessions were conducted in the control group classes against 9.5 in the treatment group classes. Throughout the paper, we estimate intention to treat (ITT) effects of assigning a class to the treatment. Given that less than 10% of the control group classes received the treatment, while almost 100% of treatment group classes received it, this ITT effect “almost” estimates the effect of delivering the treatment in a class.

\vskip\medskipamount Panel A also shows that 4.8% of eligible students in the control group attended at least one session, against 84.9% in the treatment group. Some eligible students did not attend any group session, either because their parents refused that they participate, or forgot to send back the document they had to sign to authorize their child's participation. Table (ref) compares the characteristics of the “takers”, eligible students in the treatment group that attended at least one session, to those of the “non takers” that did not attend any session. The main difference between the two groups is that the takers are less disruptive at baseline. On average, eligible students attended 0.4 sessions in the control group, against 7.4 in the treatment group. This number is 8% lower than $9.5\times 0.849=8.1$, the number we would have observed if students attending at least one session had attended all the sessions conducted in their class. This small difference is due to the fact that those students sometimes miss school on a workshop day, but school absenteeism does not seem to reduce students' exposure to the program very much. Finally, Panel A shows that the fidelity with the program's assignment rules was very high: in the treatment group, only 1% of ineligible students attended at least one session.

\vskip\medskipamount Panel B of the table shows that compliance with randomization was lower for the parents' than for the students' workshops: 53.5% of eligible parents in the treatment group attended at least one session, and eligible parents attend on average 1.0 sessions out of 3.

table[table omitted — 2,357 chars of source]

Internal validity

Balancing checks

We test for baseline differences between the treatment and control groups by estimating Regression (ref) with student- and teacher-level baseline measures as the dependent variables. First, Table (ref) compares eligible students in the treatment and control groups on 29 baseline characteristics. Only two differences are significant at the 10% level: treatment group students are more disruptive as per enumerators' ratings, and they are more likely not to be nominated as a friend by any other student in the class. Only the first of those two differences is significant at the 5% level. Second, Table (ref) compares ineligible students in the treatment and control groups on the same 29 baseline characteristics. Four differences are significant at the 10% level, one of which is also significant at the 5% level. Treatment group students have slightly worse social contact, attention and focus, activity level, and disruptiveness TOCA scores. Table (ref) compares teachers in the treatment and control groups on 12 characteristics. Only one difference is significant at the 5% level. Finally, Table (ref) compares six class-level characteristics in the treatment and control groups. Three differences are significant at the 10% level, one of which is significant at the 5% level. Treated classes are more disruptive than control ones according to teachers and enumerators, and have higher decibel levels.

\vskip\medskipamount Overall, we conduct 76 balancing checks in Tables (ref), (ref), (ref), and (ref). We find 10 significant differences between the treatment and control groups at the 10% level, four significant differences at the 5% level, and no significant difference at the 1% level.

Attrition

In this section, we document the percentage of students in our sample for which endline measures are not available, and the most common reasons for such attrition. We also show that the treatment and the control groups do not present differential levels of attrition, and that the characteristics of treatment and control group students for which endline measures are available are still balanced.

\vskip\medskipamount Table (ref) considers attrition among eligible students. Column (1) shows the levels of attrition in the control group. Endline measures collected by the enumerators are missing for 25.2% of students. For 5.9% of them this is because they have left the class between baseline and endline, for instance because their parents have moved to a different neighborhood. For the most part, the remaining 19.3% are students who were absent on the day when the enumerators visited the class.\footnote{There are also a couple of classes that enumerators could not visit at endline, because the school principal did not want to sacrifice again a half day of instruction for the purpose of the study.} The teacher's endline disruptiveness rating is missing for 23.2% of students. Again, for some of them this is because they have left the class at endline. But for the majority of students, this is because their teachers refused to rate students' disruptiveness, or only rated, say, the first half of the class and then stopped because they thought the task was too time-consuming. Column (2) of Table (ref) shows tests of differential attrition between the treatment and control groups, conducted by estimating Regression (ref) with measures of attrition as the dependent variables. Attrition does not seem differential: of the five measures we consider, only one is significantly different between the treatment and control groups at the 10% level.

\vskip\medskipamount Table (ref) considers attrition among ineligible students. Columns (1) and (2) respectively show the levels of attrition in the control group, as well as tests for differential attrition between the treatment and control groups. The attrition levels in the control group are similar to those observed among eligible students. Here again, attrition is not differential: of the five measures we consider, only one is significantly different between the treatment and control groups at the 10% level.

\vskip\medskipamount Finally, we conduct balancing checks again, among the students whose endline measures are available. Table (ref) (resp. Table (ref)) considers the same 29 baseline characteristics as in Table (ref), and compares their mean in the treatment and control groups, among the eligible students for which enumerators' endline measures (resp. the teacher's endline disruptiveness rating) are (resp. is) available. As in Table (ref), few differences are significant. Table (ref) repeats the same exercise, among ineligible students for which enumerators' endline measures are available. Again, few differences are significant. Finally, Table (ref) compares ineligible students for which the teacher's endline disruptiveness rating is available in the treatment and control groups. More differences are significant, but most become insignificant once p-values are adjusted for multiple testing. Overall, the post-attrition treatment and control group students whose outcomes are compared in Section (ref) seem to have balanced baseline characteristics.

\vskip\medskipamount Turning to class-level attrition, while we have teachers' and enumerators' ratings of classes' disruptiveness for more than 90% of classes in our sample, we have some differential attrition for teachers' questionnaires: none is missing in the control group, while 8% are missing in the treatment group, and the difference is statistically significant. In Table (ref), we conduct again the balancing checks on the baseline class-level measures in Table (ref).\footnote{Table (ref) was not pre-specified in our PAP, because we had not anticipated the possibility of differential attrition for the class-level measures.} For measures made by teachers, we restrict the sample to classes for which all class-level endline teacher measures are available, while for measures made by enumerators we restrict the sample to classes for which all class-level endline enumerators measures are available. As in Table (ref), three differences are significant at the 10% level, but none is significant at the 5% level.

Estimation methods

In this section, we discuss the methods we use to estimate the effect of the treatment. For our student-level outcomes, we estimate the following regression:

equation[equation omitted — 93 chars of source]

where $Y_{ijk}$ is the outcome of student $i$ in class $j$ and lottery $k$, the $\gamma_k$s are lottery fixed effects, $X_{ijk}$ denotes student-level baseline variables used as statistical controls, and $D_{jk}$ is an indicator variable equal to 1 if class $j$ in lottery $k$ was assigned to the treatment group. $\widehat{\beta}$ estimates the ITT effect of being assigned to the treatment on the outcome. As in Regression (ref), we cluster the standard errors at the lottery level. To select the controls, we follow belloni2014high. We run a Lasso regression of the outcome on all the student-level baseline variables in Table (ref), and we pick the variables selected by the Lasso.\footnote{In a randomized experiment, the treatment is by construction uncorrelated with the controls, so it is not necessary to run a Lasso regression of the treatment on the controls.}

\vskip\medskipamount For all the class-level outcomes, we estimate the following regression:

equation[equation omitted — 86 chars of source]

where $Y_{jk}$ is the outcome of class $j$ in lottery $k$, $Z_{jk}$ denotes class-level baseline variables used as statistical controls, and $D_{jk}$ is the treatment indicator. The regression is weighted by propensity score weights, and as in Regression (ref), we cluster the standard errors at the lottery level. To select the controls, we follow again belloni2014high, and we run a Lasso regression of the outcome on the class average of all the student-level baseline variables in Table (ref), and all the class-level baseline variables in Tables (ref) and (ref), and we pick the variables selected by the Lasso.

\vskip\medskipamount To account for multiple testing, we follow the same approach as finkelstein2010. First, we group related outcomes into hypothesis. For instance, students' happiness, self-esteem, and self-control scores are grouped together into an “emotional stability” hypothesis. Then, for each outcome, we report both the unadjusted p-value of the estimated effect, and the adjusted p-value controlling the FDR within the hypothesis the outcome belongs to. Each panel in Tables (ref), (ref), and (ref) corresponds to a set of related outcomes grouped into an hypothesis. Finally, for each hypothesis we also report the effect of the treatment on a weighted average of the outcomes in that hypothesis, using the weights proposed in anderson2008multiple. We refer to the effect of the treatment on this weighted average as the standardized treatment effect.

Treatment Effects

Effects on eligible students

Table (ref) below shows the effect of the SFL workshops on eligible students' outcomes.

table[table omitted — 2,889 chars of source]

Panel A shows that the SFL workshops do not have large effects on eligible students' emotional stability. The average school happiness score is $0.123\sigma$ higher in the treatment than in the control group, but this difference is not very significant (p-value=0.101), and becomes insignificant after adjusting for multiple testing. The average self-esteem score is $0.106\sigma$ lower in the treatment group, but this difference is insignificant even before adjusting for multiple testing (p-value=0.176). The average self-control score is very close in the treatment and control groups. Finally, the average standardized score is also very close in the treatment and control groups.

\vskip\medskipamount Panel B shows that SFL does not have a large effect on eligible students' disruptiveness. At endline, the average teachers' disruptiveness rating is $0.1\sigma$ higher in the treatment than in the control group. This difference is not statistically significant at conventional levels, but based on its estimated standard error, we can rule out at the 5% level that SFL reduces teachers' disruptiveness ratings by more than $0.1\sigma$. This is around 1/5 of the treatment effect on students' disruptiveness found by payton2008 in their meta-analysis of selected SEL programs. Enumerators' disruptiveness ratings also do not significantly differ in the treatment and control groups.

\vskip\medskipamount Panel C shows that SFL also does not have large effects on the academic outcomes of eligible students. For instance, students' Spanish and mathematics scores are very close in the two groups. We can reject at the 5% level that SFL increases eligible students' Spanish and mathematics scores by more than $0.086\sigma$ and $0.151\sigma$, respectively. Again, these effects are much smaller than those found in the meta-analysis by payton2008.

\vskip\medskipamount Finally, Panel D shows that SFL does not have large effects on eligible students' friendship ties. The proportion of students not nominated as a friend by any other student in the class is 2.8 percentage points lower in the treatment than in the control group, but this difference is insignificant.

\vskip\medskipamount Overall, we do not find evidence of a positive effect of SFL on any of the dimensions we consider, and we can also rule out much smaller effects than those previously found for similar programs.

Effects on ineligible students

In this section, we explore whether the SFL workshops have spillover effects on ineligible students. Panel A of Table (ref) below shows that these workshops do not generate strong spillover effects on the emotional stability of ineligible students. The average school happiness, self-control, and self-esteem scores are very close and do not significantly differ in the treatment and control groups.

table[table omitted — 2,893 chars of source]

Panel B suggests that the SFL workshops may generate negative spillover effects on ineligible students' disruptiveness. At endline, the average of teachers' disruptiveness ratings is $0.208\sigma$ higher in the treatment than in the control group. This difference is significant (p-value=0.05), but becomes marginally insignificant after adjusting for multiple testing (adjusted p-value=0.101).

\vskip\medskipamount Then, Panel C shows that the SFL workshops do not have large spillover effects on ineligible students' academic outcomes. Finally, Panel D shows that SFL improves the integration of ineligible students in the class network. The proportion of students not nominated as a friend by any other student in the class is 3.5 percentage points lower in the treatment than in the control group, a 17.8% reduction in the fraction of ineligible students who have no friends. This difference is significant (p-value=0.008), and it remains significant after accounting for multiple testing (adjusted p-value=0.033). Similarly, ineligible students are nominated as friends by 9.1% of their classmates in the treatment group, against 8.7% in the control group, but this difference is not significant. The treatment does not significantly alter the academic ability and disruptiveness of ineligible students' friends. Finally, the average standardized score constructed from these four outcomes is significantly higher in the treatment than in the control group (p-value=0.076).

Effects on the classroom environment

table[table omitted — 1,807 chars of source]

In this section, we study how the SFL workshops affect different measures of classrooms' environment at endline. Table (ref) above shows that SFL worsens teachers' and enumerators' disruptiveness ratings of the classes. Those ratings are based on teachers' and enumerators' agreement with statements like “There is noise and disorder in this class,” or “The teacher has to wait for a long time before students calm down and he/she can start teaching”. According to teachers, treated classes are $0.232\sigma$ more disruptive than control ones. This difference is statistically significant before adjusting for multiple testing (p-value=0.091), but it becomes insignificant after adjusting for it (adjusted p-value=0.226). According to enumerators, treated classes are $0.389\sigma$ more disruptive. This difference is statistically significant before and after adjusting for multiple testing (p-value=0.009, adjusted p-value=0.043). Enumerators do not know if the class they observe has been treated or not, contrary to teachers. The fact that they also find that treated classes are more disruptive suggests that teachers' worse perception of the treatment-group classes is not a mere placebo effect. Table (ref) in the Appendix shows that treated and control classes are imbalanced on these two measures at baseline, so we reestimate these two effects controlling for these two measures.\footnote{In the estimation of the treatment effect on teachers' ratings, the Lasso selects teachers' baseline ratings as a control, but it does not select enumerators' ratings. In the estimation of the treatment effect on enumerators' ratings, the Lasso does not select any control.} The estimated treatment effects on teachers' and enumerators' ratings are now respectively equal to $0.247\sigma$ (p-value=0.084) and $0.282\sigma$ (p-value=0.066), so the treatment effects on these two measures do not seem due to imbalances already existing at baseline.

\vskip\medskipamount It may be surprising that the treatment significantly worsens enumerators' ratings of classes' overall disruptiveness, without affecting their ratings of eligible and ineligible students' disruptiveness, as shown in Panel B of Tables (ref) and (ref). While the limited amount of time they spend in each classroom may be enough for them to observe that there is more disorder in the treated classes, it may not be sufficient for them to pinpoint the students responsible for that disorder.

\vskip\medskipamount Table (ref) also shows that treated classes have higher levels of bullying, that their lectures start 1.2 more minutes after the scheduled time than in control classes, and that they have higher levels of decibels. Even though these results are not statistically significant, they go in the same direction as the results on the disruptiveness measures.

\vskip\medskipamount Finally, the average standardized score constructed from the five outcomes in Table (ref) is $0.424\sigma$ higher in the treatment than in the control group. This difference is highly significant (p-value=0.001), and it remains highly significant even accounting for the fact that in Tables (ref), (ref), and (ref) we estimate the effect of the treatment on nine standardized scores (adjusted p-value=0.009). Therefore, we can conclude that SFL significantly worsens the studying conditions in treated classes.

Robustness checks

As a robustness check, we reestimate all the regressions in Tables (ref), (ref), and (ref) without controls. The results of that exercise can be found in Tables (ref), (ref), and (ref). Results with and without controls are pretty similar, except that the effects on ineligible students' friendships are no longer significant without controls. In our PAP, we had indicated that as a further robustness check, we would recompute all the unadjusted p-values in Tables (ref), (ref), and (ref) using randomization inference. Doing so does not change our main findings so the results of that exercise are not reported here but are available upon request.

comment\subsection{Heterogeneous treatment effects} In our PAP we had indicated that to investigate treatment effect heterogeneity, we would reestimate the treatment effect in various subgroups of students (e.g.: boys and girls). However, when estimating the treatment effect in many subgroups, one may spuriously find heterogeneous treatment effects due to type 1 error if p-values are not adjusted for multiple testing, while one may fail to detect truly heterogeneous effects due to type 2 error if p-values are adjusted. Instead, we prefer to use one of the machine-learning based methods that has been proposed since we wrote our PAP. \vskip\medskipamount Specifically, we use the method proposed by chernozhukov2018. Omitting a few technical details,\footnote{For instance, in step 4 below, the treatment has to be demeaned, and the regression has to be weighted. See chernozhukov2018 for a comprehensive description of the method.} the method amounts to repeating the following steps, say 100 times: \begin{enumerate} • Randomly split the sample into a training and a validation sample. • Train a machine-learning model to predict the outcome of the control-group training-sample observations, based on some baseline covariates. Then, train the same machine-learning model to predict the outcome of the treatment-group training-sample observations. • Use those two models to predict the treatment effect of the validation-sample observations, and divide the validation sample into, say, quartiles of the predicted treatment effect. • Regress the outcome of validation-sample observations on their predicted outcome without the treatment, their predicted treatment effect, and on indicators of predicted-treatment-effect quartiles interacted with the treatment. Let $\widehat{\theta}$ denote the difference between the coefficients of the fourth and first quartiles interacted with the treatment. \end{enumerate} To estimate the amount of treatment effect heterogeneity, chernozhukov2018 show that one can use $\text{med}(\widehat{\theta})$, the median of $\widehat{\theta}$ across the 100 replications. To compute the p-value of this estimator, the authors show that one can use the median p-value of $\widehat{\theta}$ multiplied by two. \vskip\medskipamount We use this method to investigate treatment effect heterogeneity along seven students' baseline characteristics: their gender; the social security score of their family; their mother's education; their average Spanish and mathematics score; the average of their authority acceptance, attention and focus, activity levels, and disruptiveness first-grade TOCA scores; their school happiness score; the percentage of their classmates that nominate them as a friend. We use elastic net regressions, as this is the model that performs the best in the application in chernozhukov2018. Our regressions include the seven variables listed above, their square, and the 42 products between the variables. \vskip\medskipamount In Table (ref) below, we investigate treatment effect heterogeneity for our two main outcomes: teachers' endline disruptiveness ratings, and the average of students' endline Spanish and mathematics scores. Across the split-sample replications, the median difference between the treatment effect of eligible students predicted to be in the top and bottom quartiles of the treatment effect by the elastic net is equal to $0.459\sigma$ for teachers' ratings of disruptiveness, and $0.320\sigma$ for students' Spanish and mathematics scores. These differences are both insignificant: their p-values are respectively equal to $0.206$ and $0.468$. For ineligible students, these median differences are also both insignificant. Overall, we do not find any evidence of heterogeneous treatment effects. \begin{table}[H] \begin{threeparttable} \caption{Heterogeneous treatment effects} { \begin{tabular}{lcc} \toprule\toprule &{$me(\widehat{\theta})$}&{Unadj. P}\tabularnewline &{(1)}&{(2)} \tabularnewline \midrule\multicolumn{3}{c}{Panel A: eligible students}\tabularnewline\midrule Teachers' ratings of students' disruptiveness at endline & 0.459 & 0.206 \tabularnewline Average of students' Spanish and mathematics test scores at endline & 0.320 & 0.468 \tabularnewline \midrule\multicolumn{3}{c}{Panel B: ineligible students}\tabularnewline\midrule Teachers' ratings of students' disruptiveness at endline & 0.163 & 0.435 \tabularnewline Average of students' Spanish and mathematics test scores at endline & 0.121 & 0.552 \tabularnewline \bottomrule \end{tabular} \begin{tablenotes}[para,flushleft] • {\it Notes:} This table investigates treatment effect heterogeneity for teachers' disruptiveness ratings and for students' Spanish and mathematics tests scores, using a method proposed by chernozhukov2018. Column (1) reports the median of the difference between the treatment effect of students predicted to be in the highest and lowest quartiles of treatment effect according to elastic net regressions, across 100 split-sample replications. Column (2) reports the p-value of that median. Students' baseline characteristics used in the elastic net regressions are: their gender; the social security score of their family; their mother's education; their average Spanish and mathematics score; the average of their disruptiveness first-grade TOCA scores; their school happiness score; the percentage of their classmates that nominate them as a friend. \end{tablenotes} } \end{threeparttable} \end{table}

Interpretation and exploratory analysis

Table (ref) shows that SFL does not have positive effects on eligible students' emotional stability, disruptiveness, and academic ability. This is at odds with an extensive literature, that has shown that selected SEL programs usually produce large positive effects on these dimensions. In a meta-analysis of 80 selected SEL interventions, payton2008 find that they reduce conduct problems by $0.47\sigma$, and respectively improve emotional stability and academic performance by $0.50$ and $0.43\sigma$. Based on our estimates, we can reject effects much smaller than those found in payton2008.

\vskip\medskipamount There are several other meta-analyses of SEL interventions that have been peer-reviewed, unlike payton2008, and that are more recent. However, they either focus on universal interventions delivered to the whole class rather than to a selected group of students (see durlak2011impact, sklad2012effectiveness, wigelsworth2016impact, taylor2017promoting, and corcoran2018effective), or they include both universal and selected interventions but do not report effects separately for both types of interventions (see dymnicki2012adolescent). To our knowledge, payton2008 is the only meta-analysis reporting effects separately for selected SEL interventions comparable to SFL, which is why we focus on that meta-analysis. In any case, those six other meta-analyses also find pretty large effects, even though they are slightly lower than those in payton2008. The effects they find on conduct problems range from -0.14 to -0.47$\sigma$, with an average equal to -0.25$\sigma$. Similarly, effects on emotional stability range from 0.23 to 0.74$\sigma$, and the average is 0.55$\sigma$. Finally, effects on academic performance range from 0.26 to 0.53$\sigma$, and the average is 0.28$\sigma$. Therefore, we can still reject effects substantially smaller than the average effects in those meta-analyses. Overall, our results are at odds with a very substantial literature that has studied SEL programs.

\vskip\medskipamount To understand this discrepancy, in Table (ref) below we compare SFL to the selected SEL interventions reviewed in payton2008, to assess if SFL differs from those interventions in any striking way that could account for its lower effect. Many features of the interventions reviewed in payton2008 are readily available from Table 7 therein. Other features that seemed important to us are not reported in the paper, so we reviewed a random sample of 25 of the meta-analysis's papers, and manually collected those features. They appear in italic in Table (ref).

table[table omitted — 2,972 chars of source]

SFL's intensity is comparable to that of meta-analysis's interventions

Panel A of Table (ref) shows that SFL's intensity is similar to that of the meta-analysis's interventions. The median number of sessions across those interventions is slightly higher than SFL's number of sessions (12 versus 10), but their sessions are typically shorter (50 versus 120 minutes). The number of students per workshop is comparable (a median of 6 in the meta-analysis, versus 7.2 on average in our sample). Their median duration is the same as SFL's (10 weeks). 59% of those interventions only include sessions with students, while 41% also include a parental training, like SFL. Only seven of the papers we reviewed give the number of parental sessions, but among those the median number of sessions (14) is higher than in SFL (three parental sessions). Only three of the papers we reviewed mention parents' attendance, but among those the median attendance (49%) is comparable to that in SFL (34%, see Table (ref)). Note also that payton2008 do not mention that the presence and intensity of a parental training is correlated with larger program effects. In all those interventions, selected students are pulled-out of their class during the class day, as in the SFL intervention.\footnote{The three SFL teams we interviewed said that schools' cooperation is usually very good, and that they do not have issues scheduling and delivering the sessions.}

\vskip\medskipamount Panel B of Table (ref) shows that SFL uses similar criteria as the programs reviewed by payton2008 to determine which students are eligible. Like SFL, 69% of those interventions treat primary school students. 48% target students with conduct problems, 23% target students with emotional distress, and the remaining interventions target students with a combination of problems. 73% target low SES students, like SFL.

Our study design is comparable to that of metanalysis's studies

Our study design is also comparable to that of the meta-analysis's studies. Panel C of Table (ref) shows that the treatment was randomly assigned in 80% of those studies. Many of the published studies appeared in high-impact-factor peer-reviewed journals (median impact factor=4.01).\footnote{85% of the 80 studies reviewed by payton2008 were published in peer-reviewed journals.} Most of their outcome measures are teacher, enumerator, and student ratings, often made using validated psychometric scales, as in our study. We measured our outcomes three weeks after the end of the SFL intervention, while in the reviewed interventions, the median number of weeks between the end of the intervention and endline data collection is equal to one.

\vskip\medskipamount Another methodological concern is that our control group may have benefited from the treatment, as we have some schools that have both treated and control classes, and treated students may interact with students from control classes in their school. To assess if this is a serious concern, we estimate SFL's effect in schools where only one class was included in our experiment. In this subsample, which still has 114 classes, we find that teachers' ratings of eligible students' disruptiveness is 0.2$\sigma$ higher in the treatment than in the control group (p-value=0.12), and we can rule out at the 5% level that SFL reduces eligible students' disruptiveness by more than 0.06$\sigma$. Results are similar when we consider other outcomes, such as students' test scores. Overall, control-group contamination seems unlikely to account for SFL's lack of effect.

Students receiving SFL may be harder to treat than those in the meta-analysis's interventions.

Panel D of Table (ref) shows that 85% of the interventions in payton2008 take place in the US, and all take place in high-income countries. SFL takes place in Chile, and may then be faced with a harder-to-treat population than those interventions. For instance, recent epidemiological studies show that the prevalence rate of ADHD, a disorder correlated with conduct problems, is equal to 15.5% among primary school children in Chile (see de2013epidemiology), against 6.8% in the US (see visser2014trends), 3.5 to 5.6% in France (see lecendreux2011prevalence) or 3% in Italy (see bianchini2013prevalence). Similarly, surveys indicate that domestic violence, a cause of conduct disorder problems in children (see carrell2010externalities), is more prevalent in Chile than in the US. 4.3% of Chilean women report having been physically assaulted by their partner over the previous year (see ministerio2017), against 1.3% in the US (see tjaden2000). Then, disruptive students may suffer from more severe problems in Chile than in high-income countries and may be harder to treat. This could explain why SFL produces lower effects than SEL programs in high-income countries.

\vskip\medskipamount To test this hypothesis, we start by assessing whether SFL's effect is stronger for less disruptive students. Specifically, we look at the effect of SFL for eligible students with a TOCA score below the median. If anything, we find slightly negative effects: the program increases their teachers' disruptiveness ratings by $0.18\sigma$, but this effect is marginally significant (p-value=0.09).

\vskip\medskipamount Then, as primary school students with serious behavioral problems are more present in Chile than in high-income countries, each SFL workshop is more likely to comprise some very disruptive students than an SEL workshop in a high-income country, and the presence of those hard-to-treat students may lower the workshop's effectiveness for every student, including the less disruptive ones. To investigate that possibility, we computed the 90th percentile\footnote{The choice of the 90th percentile was guided by the fact that $0.9^7=48\%$, so assuming that students' disruptiveness levels are independent within a class and that all classes have 7 eligible students, 52% of classes should have at least one student above that percentile. In practice, the proportion of classes that have at least one eligible student above that percentile is slightly lower (46%), but still close to 50%.} of the average of the authority acceptance, attention and focus, activity levels, and overall disruptiveness TOCA scores among eligible students, and estimated SFL's effects in the 79 classes that have at least one very disruptive eligible student above that threshold. Those classes have 123 very disruptive eligible students, 534 other eligible students, and 2,064 ineligible students. In Table (ref) below, we estimate the effects separately for each group of students, focusing on disruptiveness ratings and test scores, and on the friendship nominations received by very disruptive eligible students. Unadjusted p-values and p-values controlling the False Discovery Rate (FDR) (see benjamini1995controlling) across all the tests in the table are presented.

\vskip\medskipamount First, Panel A of the table shows that the program does not have any statistically significant effect on the disruptiveness and test scores of very disruptive eligible students, but increases by 50% the percentage of their classmates who nominate them as friends (unadjusted p-value=0.042, adjusted p-value=0.102). Second, in Panel B we estimate SFL's effects among the other eligible students. The program increases their teachers' disruptiveness ratings by 0.496$\sigma$ (unadjusted p-value=0.0006, adjusted p-value=0.0102), may reduce their Spanish scores by 0.201$\sigma$ (unadjusted p-value=0.033, adjusted p-value=0.112), does not have a significant effect on their enumerators' disruptiveness ratings and maths scores, and doubles the proportion that nominate at least one very disruptive student as a friend (unadjusted p-value=0.012, adjusted p-value=0.051). Very disruptive eligible students may then have a negative influence on other eligible students, which could explain the negative effects the program has on them. Third, in Panel C we estimate that the program increases teachers' and enumerators' disruptiveness ratings of ineligible students, respectively by 0.477$\sigma$ (unadjusted p-value=0.008, adjusted p-value=0.045) and 0.137$\sigma$ (unadjusted p-value=0.083, adjusted p-value=0.176). On the other hand, the program does not have a significant effect on the test scores of those students and on the proportion of them who nominate a very disruptive student as a friend. The mechanism whereby the program makes ineligible students more disruptive may be a contagion effect: eligible students become more disruptive, and ineligible students imitate them. Finally, in Panel D we estimate that the program increases teachers' and enumerators' overall disruptiveness ratings of the classes, respectively by 0.669$\sigma$ (unadjusted p-value=0.005, adjusted p-value=0.043) and 0.516$\sigma$ (unadjusted p-value=0.035, adjusted p-value=0.099). The regressions in the table are estimated with the controls selected by the Lasso. Treatment effects are similar when those controls are dropped (see Appendix Table (ref)), and when the few covariates that are unabalanced at baseline in the relevant subsample are added as controls (see Appendix Table (ref)).

\vskip\medskipamount Classes with at least one very disruptive eligible student have slightly more eligible students than classes that do not have any (8.3 versus 6.2). This difference is not very large, but we still checked if we also find negative effects of the program in the subsample of classes that have more eligible students than the median. The answer is negative, so it does not seem that the negative effects we find in classes with at least one very disruptive eligible student are mediated by the slightly higher number of eligible students in those classes.

\vskip\medskipamount Overall, we find suggestive evidence that SFL's effectiveness is hampered by the presence of very disruptive students, who may be less present in the other contexts where SEL programs have been shown to work. We still do not find statistically significant effects of SFL in the 93 classes that do not have any very disruptive student. This may be because the effects we can reject in this subsample are too large, though we can for instance still reject at the 5% level an effect larger than 0.13$\sigma$ on Spanish test scores. Another potential explanation is that even those classes may still have some students that are more disruptive than the typical students benefiting from selected SEL programs in the US.

table[table omitted — 2,976 chars of source]

SFL's delivery is less monitored than the meta-analysis interventions'.

\vskip\medskipamount Panel E of Table (ref) shows that SFL strikingly differs from the meta-analysis's programs in terms of delivery. All of the meta-analysis's interventions are demonstration programs, mounted by researchers for research purposes. 43% of the interventions are entirely or partly delivered by the researchers, 22% are delivered by school staff trained and supervised by the researchers, and 35% are delivered by other personnel (most often psychologists) hired, trained, and supervised by the researchers. 69% of the studies where the intervention was not entirely delivered by the researchers mention the frequency at which the researchers monitored the delivery personnel, for instance by attending sessions, or by reviewing video- or audio-recorded sessions. The median is a weekly monitoring. Researchers' involvement is very high in the studies reviewed by payton2008, but another meta-analysis of universal SEL interventions suggests that they can produce large effects without researchers' involvement. wigelsworth2016impact review 25 interventions implemented without researchers' involvement and find large effects: -0.15$\sigma$ for conduct problems, +0.47$\sigma$ for emotional stability, and +0.22$\sigma$ for academic performance. However, looking at a random sample of 10 of those 25 studies, it appears that in 6 of the 8 studies where monitoring was discussed, monitoring was frequent and intensive, and was often conducted by an NGO promoting the program. Overall, in the majority of the SEL interventions considered in those meta-analysis, delivery is monitored frequently by a third party.

\vskip\medskipamount JUNAEB provides SFL implementers with a detailed manual describing the content of each of the workshop's session, and the municipal teams we interviewed said they follow this manual. SFL employees also attend “good practices” meetings every six months, during which they share with other teams what seems to work in their sessions. However, JUNAEB does not systematically and frequently monitor each team's delivery. Of the three teams we interviewed, only one had a workshop observed over the last two years.\footnote{SFL employees also do not have monetary or non-monetary incentives tied to the quality of their workshops.} Then, SFL's lack of effect may be due to the lack of a frequent and intensive third-party monitoring of the workshops, unlike what is happening in the studies reviewed by payton2008 and wigelsworth2016impact. Without sufficient monitoring, teams may not implement the program with high-enough fidelity. Unfortunately, beyond the striking difference between SFL and the reviewed interventions on that dimension, we cannot further support that conjecture by testing whether the treatment effect is larger for teams that are monitored more often or that implement SFL with higher fidelity. We do not observe the frequency at which each municipal team in our study is monitored by JUNAEB, and in any case our discussions with the teams and JUNAEB officials suggest monitoring is weak in every town. Similarly, we do not have a measure of implementers' fidelity and we do not observe implementers' number of years of experience into the program, which could be a proxy for fidelity.

Conclusion

We explore the effects of “Skills for life” (SFL), a nationwide school-based SEL program for disruptive second graders in Chile. Eligibility to the program is based on first-grade teachers' ratings of students' disruptiveness, and SFL workshops consist in 10 two-hours sessions during which psychologists help students recognize and express their emotions, and teach them techniques to improve their behavior. We randomly assigned 172 classes to either receive SFL in the first or in the second semester of the 2015 school year, and we measured outcomes between the two semesters. Eligible students in treated classes see no improvement in their emotional stability, disruptiveness, and test scores. This is at odds with a large literature that has found large effects of SEL programs (see payton2008, durlak2011impact, dymnicki2012adolescent, sklad2012effectiveness, wigelsworth2016impact, taylor2017promoting, and corcoran2018effective for recent meta-analyses).

\vskip\medskipamount To understand this discrepancy, we investigate the differences between SFL and the programs studied in the literature. First, we find that SFL is not less intensive than those other programs. Second, its population may be harder to treat: all the programs studied in the literature take place in high-income countries, where the prevalence of ADHD, a disorder correlated with conduct problems, is much lower than in Chile. Accordingly, each SFL workshop is more likely to comprise one or two very disruptive students than an SEL workshop in a high-income country, and the presence of those hard-to-treat students may lower the workshop's effectiveness. We actually find evidence that SFL may increase students' disruptiveness in classes that have at least one very disruptive eligible student. The mechanism seems to be that SFL increases the friendships between very disruptive and other eligible students. Then, very disruptive students may have a negative influence on those other eligible students. To remediate this, SFL could exclude very disruptive students from its workshops, and offer them another type of treatment, for instance one-on-one sessions with a psychologist. Third, the literature has only considered small-scale programs mounted by researchers or NGOs, and either delivered by the researchers or NGO personnel, or by personnel closely monitored by them. On the other hand, SFL is a governmental program, and the government does not monitor the workshops' content and quality. Most of the interventions reviewed in the literature are implemented at a very small scale, in a handful of schools: the median number of treated students in the interventions reviewed by payton2008 is equal to 36. On the other hand, SFL treats around 8,500 students per year, in thousands of schools. Monitoring SFL as intensively as the small-scale interventions in payton2008 would probably be costly, but our results suggests this may be worth trying.

center[center omitted — 49 chars of source]