Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
52,582 characters · 24 sections · 21 citation commands
Skills to not fall behind in school
In a recent report by the World Economic Forum world2015new there is a clear paradigm shift when it comes to skills needed by today's children, youth and adults. People are rethinking the skills considered fundamental to humans in recent decades, introducing a new skill set. The authors call this core skill set “21st Century Skills”: in total there are 16 skills splitted into three categories, (i) Foundational Literacies, (ii) Competencies, and (iii) Character Qualities. The first category includes skills that encompass more technical knowledge such as math, literacy and financial education. On the other hand, the second includes skills such as creativity, communication, and the third includes skills such as curiosity, leadership, and persistence. It is important to point out that in the report itself, the authors classify the 10 skills belonging to the last two categories as part of the set of Social and Emotional skills, i.e. skills related to the way people interact to the outside world and to themselves. Figure (ref) was extracted from world2015new and exposes the 16 skills for the 21st century:
\footnotetext{Figure extracted from \citeonline{world2015new}.}
Just as we call the last 10 Social-Emotional skills we can call the first 6 as cognitive skills. Recent studies show findings by leading education researchers that these two types of skills take on great importance in defining people's future outcomes such as education, salaries and employment level heckman2006effects, kautz2014fostering. Cognitive skills, for example, have a major impact on income and employability while Social-Emotional skills, which account for more than 60% of fundamental 21st Century skills, have a positive impact on well-being and satisfaction and improving people's health through their lifestyle miyamoto2015skills.
Although cognitive and social-emotional skills are of great importance in people's lives in many ways, in this paper we will focus on how these skills can affect the academic progress of students who were attending school in 2012 and 2017 in a city in the countryside state of São Paulo in Brazil. We have a rich dataset containing demographic, socioeconomic, and personal characteristics of students, and we can account for academic progress delay in a way that minimizes measurement errors - we consider that students who had academic progress delay between 2012 and 2017 were those who have progressed less in grades than expected, that is, those who fell behind in school. Our results suggest that, in fact, both cognitive and social-emotional skills impact how students progress in school.
In this paper we are concerned about the academic progress delays that can be caused either by school failure or temporary dropout. Failure and school failure can be explained from many angles, including personal skills, family and socioeconomic characteristics. Regarding failure, in addition to more direct variables such as cognitive skills, leon2002reprovaccao gives evidence that variables such as income and parental education can have a large negative impact on the probability of failure. Regarding temporary dropout, some more concrete factors could be the need to enter the labor market prematurely arroyo1995educaccao, lack of motivation and expectations about future studies return eckstein1999youths or low parental education, which is a great proxy for the family's permanent income. In addition, students who have failed in the past form a group which is more vulnerable and prone to drop out soares2015fatores.
In a recent work carlos2019papel, the author showed evidences that cognitive and social-emotional skills can have an impact on dropping out and reaching high school, which are variables of school progress. Although we use a dataset very similar to that work carlos2019papel, we believe that this is an important paper to the literature because of the following factors: (i) In our work we used a larger number of social-emotional skills, which were not previously available and which are little explored in literature; (ii) Our dependent variable is a more general variable, which is academic progress delay, what helps us to deliver a broader message; (iii) We added an analysis of the odds of falling behind in school, which is natural within logistic regression framework; (iv) We use a Bayesian methodology, which among many benefits, we can highlight the fact that we will obtain the entire posterior distribution of the parameters of interest and not just a point estimate.
Firstly, it is important to know that a significant portion of our sample had educational progress delay between 2012 and 2017, as shown in Table (ref):
Secondly, given the importance that social-emotional and cognitive skills can play in shaping people's paths, especially regarding their success during and after school, we will be inclined to better understand how these skills may relate to falling behind in school, which may be the result of failure or not. In our dataset we have twelve skills that can be considered good or bad by common sense. Two of them can be considered cognitive skills - Literacy and Numeracy\footnote{Measured by language and math scores in a standardized test.} - and the other ten are social-emotional skills \footnote{We will give a better explanation in Section (ref).}: Assertiveness, Activity, Altruism, Compliance, Order, Self-Discipline, Anxiety, Depression, Aesthetics and Ideas. Two questions that will motivate us from this moment are: (i) How do these skills relate to falling behind in school? (ii) How can we compare the impacts of cognitive and social-emotional skills? In our case, what we will call academic progress delay, or to fall behind in school, is actually the age-grade distortion acquired between 2012 and 2017, i.e. if a student progressed less than five grades in five years (2012 to 2017) we say he/she fell behind or had an academic progress delay. The skills we are using in this paper were measured in 2012 according to some procedures detailed in Section (ref). In Table (ref), one can see the averages of various skills measured in standard deviations\footnote{The variables were standardized in the sample to have zero mean and unit standard deviation.}, conditioned on the variable 'Behind', which indicates students which fell behind between 2012 and 2017:
It can be seen in Table (ref) that students who fell behind between 2012 and 2017 have lower average grades in language and mathematics and differ, on average, from the group that did not fall behind with respect to various social-emotional characteristics such as depression - which leads us to believe that among our students, cognitive and social-emotional skills may also be related to whether or not a student fell behind between 2012 and 2017. Although we present a characteristic of the distribution of social-emotional and cognitive skills conditional on the variable 'Behind', we will be more concerned in this paper to estimate the distribution of the 'Behind' variable conditioned to a vector of characteristics, paying more attention to social-emotional and cognitive skills. Therefore, the descriptive analyzes done so far only serves as motivation for a more robust analysis.
The main objective of this paper is to better understand how cognitive and social-emotional skills relate to academic progress delay, which may be the result of failure or temporary dropout. For this purpose we use a dataset of approximately 1800 students who studied in the city of Sertãozinho in the countryside of the state of São Paulo and who were in elementary, middle school in 2012 and were re-interviewed in 2017.
The datasets used in this study come from two field surveys conducted by LEPES/USP ("Laboratório de Estudos e Pesquisas em Economia Social") in the city of Sertãozinho, in the countryside of the state of São Paulo, in 2012 and 2017. In both years, information was collected about the students' family and socioeconomic context, about the situations they faced in school and about cognitive and social and emotional skills. Since cognitive and social-emotional variables are key variables in our analysis, it is important to explain in detail how they were constructed.
To assess the level of cognitive development of students in 2012 and 2017, we used the result of Language and Mathematics tests prepared by the psychometrist Dr. Ricardo Primi from items available on the platform of the National Institute for Educational Studies and Research Anísio Teixeira (INEP). The test preparation methodology includes the use of the Item Response Theory (IRT), where the final grade is not the number of correct answers, but the student's proficiency level taking into account the difficulty of the items, for example. In 2012, the students, who were in the 5th and 6th grades, answered only one test of each subject. It is important to say that the students who were in the 4th grade in 2012 answered the Big Five Inventory (BFI), discussed bellow, but did not take the exam due some bureaucratic issues.
Regarding the level of social-emotional development, the researchers who designed the research chose to use a well-established social-emotional assessment scale in the literature, which is the Big Five Inventory, or BFI, developed by john1999big. Although, the BFI scale is based on the theory that a person's personality can be roughly described by big five factors, we focused in more recent measures presented in soto2009ten which are called 'facets'. We chose to work with facets instead of using the Big Five traits because they are less broad and any result achieved could be interpreted more deeply. The ten available facets and their description, according to piedmont2013revised, are:
Firstly, we should mention that a higher score in the facets (a.k.a. social-emotional skills) can not always be considered a good thing. Secondly, it is important to say that we do not use the raw scores obtained after applying the scale due to the acquiescence bias. That bias is the tendency of an individual to agree or disagree with statements about his or her attitudes regardless of those winkler1982controlling. According to valentini2017influencia, if the bias is not corrected the factor estimation may be uncorrelated with the content of the items, which would invalidate the instrument. Given that, all scores obtained for the social-emotional dimensions already discussed were corrected for acquiescence bias.
Part of the variables selected to be used in this work was selected from the 2012 dataset, part was selected from the 2017 dataset. More information about our variables:
In Appendix (ref) one can see the main descriptive statistics of the variables used in the work.
Our baseline of students will consist of the 5th and 6th grade students who were present in 2012 and had no missing values in all variables except the "Pre-K", "Kinder" and "Behind" variables, which depends on the 2017 dataset. In 2012 we have 3223 students who are distributed in 2 school grades: 5th and 6th years. Among these 3223 students, 83,71% will compose our baseline according to the rule states above. Given that, we have 2698 students in our baseline, as one can check in Table (ref):
It is important to state that our baseline is not our actual sample for this study for one major reason: we do not have data for some students in 2017, mainly because we could not reach them in the 2017 field research. Thus, we will use then the amount of 1780 students, which is 65,97% of the initial amount. Table (ref) provides more details about all this information:
The problem of losing part of the sample between two time points is called attrition. When the attrition is random, that is, when the loss is independent of the characteristics of the students, we do not have many problems. However, when we lose part of the sample in a non-random way, it may be that our results are not extendable to the population of interest, that is all students from the 5th and 6th grades of the Sertãozinho school network in the countryside of the state of São Paulo. If one check in Appendix (ref), one will see that the loss was not random: attrition was primarily related to students with less-educated mothers and those with the more problems in school.
Despite a random loss being a sufficient condition for an analysis with extendable results for the population, it is not a necessary condition in our case - I will briefly explain why. Suppose that $Y$ is a binary variable that indicates academic progress delay between 2012 and 2017 for an individual and $X$ a vector of characteristics of the same individual - draws of $X$ are observable in our data. Note that our interest is to better understand the following quantity $\mathbb{P} (Y = 1 | X = x)$, however, we can only estimate $\mathbb{P}(Y = 1 | X = x, A = 0)$ with our data, if $A$ is a variable that indicates attrition. Now imagine that a $Z$ variable vector is the only common cause of falling behind and attrition, and that $U$ and $V$ are other variables that impacts on progress delay or attrition as illustrated in the following causal DAG (Directed Acyclic Graph):
If the DAG is valid and $Z$ is a vector of variables included in $X$, then, given $X=x$, $A$ and $Y$ are conditionally independent and $\mathbb{P}(Y = 1 | X = x) = \mathbb{P}(Y = 1 | X = x, A = 0)$ pearl2009causal. It is reasonable to think that this condition is valid since our dataset is rich in demographic, school and personal characteristics of students. On the other hand, if the condition is not valid, we may think that our results will be extendable to a portion of the population, and it is still possible to speculate possible outcomes according to the differences shown in Appendix (ref). Although it is not possible to estimate $\mathbb{P}(Y = 1 | X = x)$ directly, we will use this notation from now on taking our sample as given.
The model we will use in our analysis is the Bayesian Logistic Regression model. Using that model we want to directly model the logarithm of students' odds of falling behind in school given their characteristics using a linear predictor. If $Y_i$ is a variable that indicates academic progress delay between 2012 and 2017 for student $i$, $x_i^\top = (1 ~ x_{i1} ~ ... ~ x_{ik})$ is a feature vector for that student, $\theta^\top = (\theta_{0} ~ \theta_{1} ~ ... ~ \theta_{k})$ a parameter vector that helps us to parameterize our model we define:
Since $\pi (x_i, \theta)$ is the probability of falling behind conditional on the characteristics $ x_i$ and on the $\theta$ parameters, we have that the log of the $ i $ student's odds to have a progress delay between 2012 and 2017, given $\theta$, is:
Although logistic regression models the log of the odds, we can obtain $\pi(x_i, \theta)$ and get the conditional probability of progress delay as follows:
Where $\sigma (.) $ represents the Sigmoid function (or standard logistic distribution function).
In the Bayesian framework, it is common to infer about parameters as follows: (i) we adopt an prior distribution for the parameter vector, which captures our knowledge of the quantities of interest before looking at the data; (ii) after observing the data, we update our knowledge about the parameter vector by applying the Bayes Theorem.
At first we assumed no correlation between the $\theta$ coordinates, and for each of the inputs we assumed a prior Normal distribution $N(0,1000)$, fulfilling the idea of being a weakly informative prior distribution. Our primary objective is to obtain the posterior distribution of the parameter vector, which is given by Equation (ref):
Given $Y_{1:n}$ is a vector and $ X_{1:n} $ is a matrix for our entire actual sample ($n$ is the sample size). The major obstacle to calculate the posterior distribution analytically is the calculation of the integral in the denominator, which is intractable. Given that, we resort to a variation of the Hamiltonian Monte Carlo (HMC) betancourt2017conceptual algorithm so we can sample from the posterior even without knowing its closed form. That variation of the algorithm is called "No U-Turn Samples" and it is implemented in "RStanarm" package \footnote{\url{https://cran.r-project.org/web/packages/rstanarm/index.html} - accessed in 11/11/2019} in R - that package is built upon "RStan". In our sampling process, we sampled 4 chains, each of which had a burn-in period equals $1000$, a thinning of $100$ and a size of $2500$ (we have $ L = 10000$ samples in total). We conducted a convergence diagnosis of MCMC which can be viewed in more details in Appendix (ref). After sampling a reasonable number of times from the posterior distribution, we can move on to the next step.
In order to understand whether the logistic regression model is a good model for our problem, we will apply a simple cross-validation procedure (one training set and one testing set) for comparing the Receiver Operating Characteristic (ROC) curves and AUC metric (Area Under Curve) between our model and a tuned Random Forest benchmark model \footnote{The model will be tuned in a cross-validation 3-Fold procedure varying the (i) number of trees in the list c(50, 100, 150, 200, 250, 300, 400, 500, 600), (ii) the number of features used in the list c(2,3,4,5,6, 7), (iii) minimum size of nodes in c(1, 3, 5, 10, 15, 20, 30), (iv) sample fraction used in bootstrap in c(.5, .6, .8, 1) and (v) replacement in sampling in c(TRUE, FALSE).}, which has high predictive power. To estimate the conditional probability of falling behind in school using bayesian logistic regression, we will use the concept of posterior predictive distribution, which is calculated as follows:
From this moment, we can interpret $ i $ as an out-of-sample individual. As long as we expect to independently sample $\theta$ from its posterior distribution, we may resort to the Law of Large Numbers to approximate the above integral by the following mean:
Where $\theta \sim p (. | Y_{1:n} = y_{1:n}, X_{1:n} = x_{1:n}) $, i.e., it is sampled from its posterior in a total $ L $ times. Regarding the methodology of comparison between two classifiers, we chose the ROC curves and the AUC metric because they do not depend on the cutoff threshold for classification and are easily interpretable: (i) the ROC curves graphically give us a tradoff between True Positive Rate and True Negative Rate of a binary classifier and (ii) the AUC metric is actually equivalent to the probability that a binary classifier will rank higher an instance of type '1' compared to an instance of type '0', chosen at random fawcett2006introduction.
The first analysis we will conduct is direct from obtaining the posterior distribution of the parameter vector. Given the property of the linearity seen in Equation (ref), we have that the percentage change due to the conditional odds of falling behind due to a $\delta$ change in $x_{ij}$, keeping all other variables constant, is given by:
It is important to note that one important implication of linearity of the predictor is that the amount $\Delta_\delta \% O(x_{ij}, \theta) $ does not depend on $ i $, but only on $ \theta_ {j} $ and on $ \delta $. In our analysis, we will consider $\delta = 1 $, which is natural for both binary variables and quantitative variables that are measured in standard deviations. Considering $\delta = 1 $, we have:
Remember that because $ \theta_ {j} $ is a random variable, $\Delta_1 \% O (x_{.j}, \theta) $ is also a random variable and that is why we will examine its distribution directly.
Our analysis regarding the importance of each of the variables in predicting progress delay in school will be analyzed in the light of a hypothesis test that we will propose inspired by esteves2019pragmatic. In the proposed framework we have three possible hypotheses for the importance of variable $j$:
Choosing one of the three options will be given by a decision problem under uncertainty. Recalling that $ \theta_ {j}$ assumes values in $ \Theta_j = \mathbb{R} $, we want our hypothesis to be as follows:
Given $\varepsilon_{1}<0$ and $\varepsilon_{2}>0$. The problem with this approach is that it is not straightforward to find $ \varepsilon_ {1} $ and $ \varepsilon_ {2} $ values that make sense. To make this task easier and the result more interpretable, first realize that $ \Delta_1 \% O (x _ {. j}, \theta) = \text {exp} \big (\theta_ {j}) - 1 $ is an increasing function in $ \theta_ {j} $ and with range equals to $ \Omega_j = (- 1, \infty) $. Equivalently, we can rewrite the hypothesis as follows:
In the way we lastly presented the hypotheses, one can realize it is easier to choose values of $\varepsilon_ {1}'$ and $ \varepsilon_ {2}' $ that make sense. To define these values, let us remember that the percentage of students who fell behind in school between 2012 and 2017 is $ p = 0.1303 $, so the odds of delay is given by $ \frac{p}{1-p} = 0.1499$. If $ p $ is added by $ d $ points (e.g. $ d = \pm 0.01 $) we get the odds to be $ \frac{p + d}{1- (p + d)} $ and the percentage change in the odds will be given by:
Given constants $d_1<0$ and $d_2=-d_1$, we define $ \varepsilon_ {1} '= \omega_{d_1} $ and $ \varepsilon_{2}' = \omega_{d_2} $. The rationale can be elucidated with the following example. Suppose $ d_1 = -0.01 $, so $ \varepsilon_{1} '= - 8.72 \% $ represents the percentage change in the odds of progress delay in school, based on our sample, due to the decreased probability of delay $ p $ by one percentage point. If $ d_2 = 0.01 $, then $ \varepsilon_{2} '= 8.92 \% $ represents the percentage change in odds of progress delay in school, based on our sample, due to the increased probability of delay $ p $ by one percentage point. In a way, the value of $ d_1 $ and hence the value of $ d_2 $ is still somewhat arbitrary, so in our analysis we will assume that $ d_1 $ and $ d_2 $ assume values in $\left \{\pm 0.01, \pm 0.02, \pm 0.03, \pm 0.04, \pm 0.05 \right \} $ to see what our decision would be like in several scenarios.
What remains to be understood is how our decision will be made after we set $ d_1 $ and $ d_2 $. If we define a constant loss function for each of the possible decision errors, i.e. there is no reason to believe that we should value differently different types of errors, our decision to $ \Delta_1 \% O (x _{. j}, \theta) $ with respect to the hypotheses $ H_0, ~ H _- $ and $ H _ + $ is given by:
We are interested in calculating the effect of a marginal increase in each of the skills on the predicted probability of school delay for an individual. That is, if $ x_{ij} = \text {skill} _ {ij}, ~ j \in \{1, ..., 12 \} $ denotes the value of $ j $-th skill for $ i $-th individual (out of sample) we want to calculate the amount:
By sampling $ L = 10000 $ from the posterior of $ \theta $, we can approximate our quantity of interest as follows:
Note that the marginal effects depends on the characteristics of each individual, so we will analyse the distribution of the marginal effects across our sample.
The first step in understanding our results is to analyze the posterior distribution of parameters related to skill variables \footnote{Complete results can be found in Appendix (ref)}. In order to make the results more interpretable and easy to understand, let us look at the marginal posterior distributions, not taking in consideration the possible relationships between the parameters:
In Table (ref) one can see the HPD (Highest Posterior Density) intervals for the posterior distributions of our parameters - we highlight the fact that all parameters of our interest had a unimodal marginal posterior distribution. An interesting point to note is that only four of the marginal distinctions of the parameters of interest had HPD intervals that did not include zero - this gives us clues as to which may be the most important variables in predicting progress delay in school. Two of the variables we have more evidence that can predict our variable of interest have a negative impact on progress delay in school (language and math scores) and two other variables we have the most evidence that can explain our variable of interest have a positive impact on progress delay in school (depression and assertiveness scores). We will put more effort to understand the importance of variables in the next sections. In Table (ref) one can see the main statistics for marginal posterior distributions:
A fundamental thing to make our results interesting is the ability of our model to fit well the data we are using. To see how well our model fits the data, we proposed a cross-validation scheme comparing the predictive power of our model with a benchmark model, which in this case will be a Random Forest model. With respect to the Bayesian logistic regression model, we use the predictive distribution of $ Y $ given $ X = x $. In relation to the Random Forest model, we used the combination 'number of trees'=200, 'number of variables per tree'=2, 'minimum node size'=1, 'replacement in sampling'=True, 'fraction sampled in bootstrap'=0.6 chose by means of a grid search in the training set (3-fold CV) in order to maximize the AUC metric. In Figure (ref) one can see a comparison between the two models:
Looking at the figure above, it is possible to see that our Bayesian logistic regression model performed better in a classification task when compared to a Random Forest tuned model, which we traditionally consider with great predictive power. We can reiterate our result by looking at Table (ref):
The great lesson we get from the results seen in this section is that our model could fit well to our data even when compared to more flexible models like a tree ensemble model, that is the random forest.
Although we get an important message by looking at the results displayed in Tables (ref) and (ref), interpretation may not be as straightforward as desired, especially when parameters assume larger magnitudes. Thus, using the methodology outlined in Section (ref), we obtain the distribution of percentage changes in odds (conditional on the $ x $ feature vector) of falling behind in school given an increase in the magnitude of a standard deviation in a skill of interest. In Table (ref) we can see the HPD intervals of the percentage change distributions in the odds of falling behind in school due to an increase in the magnitude of a standard deviation in a skill of interest:
An important thing to note is that if the HPD intervals present in Table (ref) do not contain zero, the HPD intervals in Table (ref) should not contain it either, for a simple mathematical property. Thus, the four skills that bring us the most evidence of their importance are the same as before: language score, math score, depression score and assertiveness score. In this case, for example, an increase in the math score with magnitude of one standard deviation changes the predicted conditional odds of falling behind in school at a magnitude between $ -44 \% $ and $ - 16 \% $ with a probability of $ 0.95 $. Regarding the depression score, an increase in the magnitude of one standard deviation of this score changes the predicted conditional odds of falling behind in school at a magnitude between $ 17 \% $ and $ 56 \% $ with probability of $ 0.95$. The same exercise can be done for other skills. In Table (ref), we can see some important statistics:
The first thing to note is that the odds percentage change distributions are more or less 'balanced' in the sense that the averages approach the medians and the quartiles are more or less symmetrical with respect to the median. One big message we can get from Table (ref) is that, according to the selected statistics, some social-emotional variables can be as important as cognitive skills (math and language) in defining progress delay in school, which is very interesting, given that we traditionally give more importance to the cognitive ones.
In the author's point of view, in addition to having an overview of how factors can determine educational progress delay, it is interesting to determine a rule of choice in order to test which factors really matter. In order to make the choices, let's follow the hypothesis testing framework proposed in Section (ref). As already mentioned, we performed tests for different values of $ d $, which are exposed on the upper horizontal axis of Table (ref):
In Table (ref) we can see how our decision - regarding the importance of each skill in determining educational progress delay - varies as one chooses different values of $ d $. In the table above, the symbol "0" means that, by adopting variations in percentage points for the prior probability of educational progress delay with magnitude $ | d | $, the variable of interest has no importance in predicting educational progress delay. The symbols "-" and "+" states that the variable of interest is important (negatively and positively) in predicting educational progress delay when adopting variations in percentage points for the prior probability of educational progress delay with magnitude $ | d | $. It is interesting to note that in addition to the four variables we have already mentioned, the social-emotional score of "Order" seems to be related to school delay, which is not very intuitive. However, this importance seems to be less robust when compared to the importance of the other four already mentioned variables.
Finally, we will present the results regarding the methodology presented in Section (ref). The marginal effects we calculate give us the rate of change in conditional probability of school delay for an infinitesimal variation in the $ j $ skill score, which is a quantitative variable. As we have seen, this rate of change depends on the characteristics of the individuals in our sample - in the end of the day, we will have a distribution of marginal effects in our sample. In Table (ref) one can check the HPD intervals for the marginal effects of each variable:
It can be seen in Table (ref) that the language and math scores variables and the depression and assertiveness scores continue to stand out from the other variables. Regarding the language score, for example, the result tells us that a change in $ \epsilon $ standard deviations in the score makes us expect a change that can range from $ 0 $ to $ -0.06 * \epsilon $ percentage points in conditional probability of falling behind in school, with a probability of $ 0.95 $. For the result to make sense, $\epsilon $ must be a small number, but even for $ \epsilon = 1 $ we get a reasonable approximation. Regarding the assertiveness score, a variation in $ \epsilon $ standard deviations in the score makes us expect a change that can range from $ 0 $ to $ 0.06 * \epsilon $ percentage points in the conditional probability of falling behind in school, with $ 0.95 $ of probability. In Table (ref), we have some important statistics:
The results presented in Table (ref) summarize everything presented so far. We emphasize again that an interesting result obtained is that, by comparing the magnitude of the effects, social-emotional skills are as important as cognitive skills in determining educational progress delay.
The main objective of this study was to analyse the relationship between student skills and the fact of falling behind in school, since traditionally in the literature, only the relationship between the contextual variables (family and school) with school delay is evaluated - It is also important to mention the fact that It is hard to have access of a dataset containing a social-emotional assessment of students. To tackle our objective, we used a longitudinal dataset built in two field researches (2012 and 2017) in the city of Sertãozinho, in the countryside of the state of São Paulo. According to the results achieved, we can highlight the following findings, which summarize the contributions of this paper: (i) not all skills are important in determining academic progress delay, in fact, we had only four (out of a dozen) that seem to be more important, namely language and math skills (negative impact on delay) and depression and assertiveness (positive impact on delay); (ii) in terms of magnitude of impact, we can say that social-emotional skills are as important as cognitive skills in determining academic progress delay.
An additional point that was not covered during the development of this paper is that the results found potentially have a causal interpretation, although I believe that a more detailed study has to be done later. I say this because the context variables are exogenous, we got a great diversity of control variables and, in my view, the only obstacle would be the fact that social-emotional skills have an impact on the probability of falling behind in school through academic performance and did not model that relationship directly. In fact, the hypothesis that social-emotional skills cause cognitive skills is a relevant and supported hypothesis in the literature cunha2008formulating, cunha2010estimating - the opposite way is still more questionable. During the experimental phase of this work, I estimated a model without math and language grades, so we would no longer have undesirable control blocking the a path between social-emotional skills and school progress delay. The results were almost identical (posterior distributions) for the social-emotional variables. Because of this, I once again highlight the potential for causal interpretation brought by the results, even if this subject should be the main topic of another paper.
To conclude, I think this work may have an influence on the debate of educational public policies in Brazil and in the world since it deals with current and important issues besides being able to change the way we look at education and what kind of skills we would like to develop in youth.
I would like to thank CAPES (Higher Education Personnel Improvement Coordination), which is financing my master's degree in statistics at IME/USP and the LEPES/USP (Laboratório de Estudos e Pesquisas em Economia Social) for the dataset used.