EconBase
← Back to paper

Machine Learning Classifiers Do Not Improve the Prediction of Academic Risk: Evidence from Australia

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

51,496 characters · 13 sections · 18 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Machine learning classifiers do not improve the prediction of academic risk: Evidence from Australia

\pretitle{

center[center omitted — 199 chars of source]

}

\preauthor{

center[center omitted — 347 chars of source]

}

\twocolumn[ \begin{@twocolumnfalse}

abstract{ \fontsize{11pt}{12pt}\selectfont Machine learning methods tend to outperform traditional statistical models at prediction. In the prediction of academic achievement, ML models have not shown substantial improvement over logistic regression. So far, these results have almost entirely focused on college achievement, due to the availability of administrative datasets, and have contained relatively small sample sizes by ML standards. In this article we apply popular machine learning models to a large dataset ($n=1.2$ million) containing primary and middle school performance on a standardized test given annually to Australian students. We show that machine learning models do not outperform logistic regression for detecting students who will perform in the `below standard' band of achievement upon sitting their next test, even in a large-$n$ setting. Keywords: Education Economics; NAPLAN; Standardized Testing. JEL Codes: I20; I29; C45; C55. \\ }

\end{@twocolumnfalse} ]

{ {\fnsymbol{footnote}}

\footnotetext[1]{{This research is supported by an Australian Government Research Training Program (RTP) Scholarship. Thanks must also be given to the Australian Curriculum, Assessment and Reporting Authority for the provision of the data utilised by this study. The authors would like to thank Foivos Diakogiannis and Airong Zhang for their helpful comments and suggestions.}}

\footnotetext[2]{Jupyter notebooks containing R codes and outputs for this paper as well as instructions on how to obtain the data set may be found at \href{https://github.com/RobGarrard/A-Machine-Learning-Approach-for-Detecting-Students-at-Risk-of-Low-Academic-Achievement}{Github.com/RobGarrard}}

}

\pagestyle{fancy} \fancyhead[HCO]{S. Cornell-Farrow and R. Garrard} \lhead \rhead

\interfootnotelinepenalty=10000

Introduction

Educational achievement is closely linked to an individual's economic well-being as well as a nation's standard of living Barro1991. In Australia, the Report of the Review to Achieve Educational Excellence in Australian Schools 2gonski highlighted the decline in student outcomes across the past twenty years when compared to other OECD countries. Not only are a large number of Australian students achieving poor outcomes, but there is a significant disparity in achievement within the same age groups and classrooms. 'One-size-fits-all' policy approaches have not been successful in bringing under-achieving students up to minimum standards. In order to support these students in reaching their full potential, it may be necessary to tailor policies and teaching strategies to individual students who are at risk of low achievement. Since schools, especially public schools, tend to face binding resource constraints, it is essential to detect `at risk' students early and with high precision.

In this setting, the goal is not to conduct causal inference on a coefficient of interest but to produce a purely predictive model that classifies students with high sensitivity and specificity. Machine learning (ML) methods tend to greatly outperform traditional statistical models Friedman2001. For example, when attempting to classify images of handwritten digits correctly, standard logistic regression achieves an accuracy of around 70%, whereas state-of-the-art neural networks obtain 99.79% accuracy Wan2013.

The application of machine learning methods to education data has been referred to as Educational Data Mining Romero2007. Its use in the prediction of academic performance has been predominantly in the higher education context Vandamme2007, Kotsiantis2012, Yadav2012, Jishan2015 largely due to the availability of administrative data sets collected by universities.\footnote{See \shortciteA{Shingari2017} for a recent review.} Interestingly, when ML models are benchmarked against logistic regression, they show no or unsubstantial improvement Kotsiantis2004, Cortez2008, HUANG2013, Gray2013. This could be because the conditional expectation function in the predictors used in these studies does not exhibit significant non-linearities, or could be that the sample sizes used are too small for the non-linearities to be learned.

In this article we exploit a large data set containing scores on the Australian National Assessment Program - Literacy and Numeracy (NAPLAN); a standardized literacy and numeracy test sat by all students in Australian schools in grades 3, 5, 7, and 9. This data set contains raw scores for all students who sat the test in the years 2013 and 2014, as well as administrative data on students' individual- and family-level characteristics. In total, the data set contains observations on 2.2 million unique students. Students are labeled into two classes, `At Standard' and `Below Standard', according to whether or not their score meets minimum achievement standards as determined by the Australian Curriculum, Assessment, and Reporting Agency (ACARA). This is done for two learning areas: literacy and numeracy. We split students into those in grade 3, for whom this would be their first time sitting NAPLAN, and students in grades 5 and above, for whom their previous achievement on NAPLAN may be used as a predictor.

We train a set of popular machine learning classifiers with standard logistic regression serving as a benchmark. We measure model performance using area under the ROC curve (AUC) and find that none of the machine learning models outperform logistic regression. To the best of our knowledge, this is the first paper to apply machine learning methods to the prediction of primary and middle school student achievement in a `big data' (large $n$) setting.

While machine learning methods have great potential for improving the analysis of economic data, the results of this study show that large data sets and state-of-the-art algorithms do not automatically imply better predictions.

Data

NAPLAN is a set of standardized literacy and numeracy tests sat by all students in Australia in grades 3, 5, 7, and 9, in both the government and non-government schooling sectors. It is described as providing “the measure through which governments, education authorities and schools can determine wh-ether or not young Australians are meeting important educational outcomes.”\footnote{See \href{http://www.nap.edu.au}{nap.edu.au}} The tests cover five learning areas known as `test domains': Reading, Writing, Spelling, Grammar and Punctuation, and Numeracy. The tests are designed to assess student performance relative to the Australian federal curriculum.

We use individual student-level data from the NAPLAN reading and numeracy tests administered in 2013 and 2014. The NAPLAN reading test involves a reading comprehension-style test on a range of texts, including imaginative, persuasive, and informative. The questions are designed to test knowledge and interpretation of English language in context. The NAPLAN numeracy tests assess students on their performance in mathematics, namely number and algebra, measurement and geometry, and statistics and probability.\footnote{Samples of these tests may be found at \url{https://www.nap.edu.au/naplan/the-tests}.}

The data set contains 2,235,804 unique student IDs, who are attending 9,250 different schools in both the public and private sectors in all states across Australia. All individuals in the data set are unique, as a student sitting NAPLAN testing in 2013 would not sit the test in 2014, and vice versa. The individual scores for each student in each domain are collected by the Australian Curriculum Assessment and Reporting Authority (ACARA), alongside student background information which is collected by schools from students' parents or carers via enrolment forms. (ref) provides a detailed summary of the variables contained in the data set.

For each testing domain in each year, ACARA determines achievement bands to classify a student's level of achievement based on what particular skills they can perform, e.g., addition of simple numbers, understanding of probability, etc.. For each year level, students in the lowest two bands of achievement are deemed to be achieving `below minimum standards'. Given the raw scores for each student available in the data, we have classified each student into their relevant achievement band following the cut-off scores published by ACARA.\footnote{\url{nap.edu.au/results-and-reports/how-to-interpret/score-equivalence-tables}} Using these bands, we have labeled each student as either `At Standard' or `Below Standard' for reading and numeracy respectively. Within this sample, approximately 16-17% of students are achieveing below standard performance in each learning domain.

In addition to each student's score in the 2013/14 NAPLAN cycle, the data also contains the score each student received in their previous testing cycle, two years earlier. This data is only present for students in grades 5 and above, as grade 3 is the first year NAPLAN is administered. In order to exploit this presumably strong predictor, we split the data set into students in grade 3, and students in grades 5 and above. We use these raw scores to construct a dummy variable for whether the student was `Previously At Standard' or `Previously Below Standard'.

Methods

Our objective is to predict whether or not a student will perform in the `Below Standard' band upon sitting their next NAPLAN.

We split the data set into two subsets: students in grade 3, for whom this is their first time sitting NAPLAN; and students in grades 5, 7, and 9, for whom their NAPLAN score in the previous cycle may be used as a predictor. Within each data set, we remove any rows that contain missing data. For each data set (grade 3, $n=345,817$; and grades 5+, $n=886,392$) we obtain a random two thirds/one third split for training and test sets respectively. We stratify the random samples such that the proportion of students not meeting minumum standards is preserved in the training and test sets respectively. We do this for each of the response variables (reading and numeracy), resulting in each classifier being estimated separately on 4 data sets (reading/numeracy, grade 3/grade 5+) and evaluated on corresponding test sets.

In what follows, let $\left(y_{i}, \mathbf{x}_{i}^{\prime}\right)_{i=1}^{n}$ be a set of $n$ iid observations of a dependent variable $y_{i} \in \{0, 1\}$ with a corresponding vector of $p$ predictors $\mathbf{x}_{i}^{\prime} \in \mathcal{X}$ in a feature space $\mathcal{X}$.

Classifiers

The task of a binary classifier is to predict whether or not an observation belongs to class 0 or 1 as a function of the predictors, $\mathbf{x}_{i}^{\prime}$. This is often accomplished by producing a predicted `probability' that the dependent variable belongs to the positive class, $\hat{y}_{i} \in [0, 1]$, and classifying the observation to that class if $\hat{y}_{i} \geq 0.5$.

A classifier's parameters are chosen such that the classifier's predictions minimize a specified loss function, such as mean square error or cross-entropy. Since we have a binary classification problem where one of the classes (`At Standard') is more common than the other (`Below Standard'), we use the weighted binary cross-entropy loss function.

multline[multline omitted — 143 chars of source]

Where each $w_{i}$ denotes a weight allowing some observations to contribute more or less to the loss function than others. Since the `At Standard' class is around 6 times more common than the `Below Standard' class, the loss function is weighted so that misclassifications of a `Below Standard' student incur around 6 times more cost. That is, $w_{i} = 6$ if the student performs below standard ($y_{i} = 1$), and $w_{i} = 1$ if the student performs at standard ($y_{i} = 0$).

Each model we consider gives a different functional form for generating $\hat{y}_{i}$ from the predictors. Unless otherwise stated, we choose the model parameters to minimize (ref). We estimate the following types of classifier.\footnote{In addition to these classifiers, other common choices include k-nearest neighbors and support vector machines. However, computation of these classifiers does not scale well with number of predictors and sample size respectively, and so were omitted from this study.}

Logistic Regression

We use plain logistic regression as a benchmark for the ML classifiers. Logistic regression models the probability of belonging to the positive class as

equation[equation omitted — 170 chars of source]

from which predicted values, $\hat{y}_{i} = \hat{\mathbb{P}}(y_{i} = 1)$, can be extracted by substituting in the predictor values and rearranging.

Elastic Net

The elastic net Zou2005 is a hybrid of the lasso Tibshirani1996 and ridge regression Hoerl1970 making it a shrinkage estimator. Shrinkage estimators attempt to exploit the bias-variance tradeoff by introducing a small amount of bias through shrinking estimated coefficients towards zero. Shrinkage estimators are especially popular in high-dimensional regression, where $p >> n$, as they can perform automatic model selection by setting many coefficients exactly to zero. The elastic net has been used in genetic microarray studies for identifying genes associated with cancer risk Zou2005, Li2010, Li2013.

The elastic net uses an identical functional form for generating predicted values as logistic regression but uses an alternative loss function for estimating the coefficients. Its loss function takes the form

equation[equation omitted — 133 chars of source]

where $\lambda$ is a tuning parameter determining the strength of the combination $\ell_{1}$ (lasso) and $\ell_{2}$ (ridge) penalty. Setting $\lambda = 0$ retrieves the unpenalized logistic regression, while setting $\lambda$ sufficiently large forces all coefficients to be zero. The parameter $\alpha$ controls the relative strength of the lasso and ridge penalties. The lasso penalty forces some of the coefficients to zero, performing model selection; while the ridge penalty forces correlated predictors to be assigned similar coefficients.

We impose that $\alpha = 0.5$, such that the lasso and ridge penalties have equal weight and select $\lambda$ through 10-fold cross-validation.\footnote{The mixing parameter, $\alpha$, may also be tuned. However, this option is not included in the glmnet package. We re-estimated the classifiers with $\alpha= 0.1$ and $\alpha = 0.9$ with no change to the results.}

Decision Tree

figure*[figure* omitted — 813 chars of source]

Classification and Regression Trees (CART) are a class of non-linear models that are tractable, flexible, and highly interpretable Breiman2017. Decision trees can be visualized in a tree diagram representing a set of `if-then' conditions on the predictors, making them easily understood and applied by humans. Further, trees can naturally handle mixed data types and require only small modification to switch between a classification tree and a regression tree.

Let $R_{1}, \dots, R_{M} \subset \mathcal{X}$ be a set of regions that partition the feature space; i.e., $R_{i} \cap R_{j} = \emptyset$ for $i\not = j$ and $\cup_{i} R_{i} = \mathcal{X}$. Let $N_{m} = \sum_{i} I(\mathbf{x}_{i} \in R_{m})$ be the number of observations that fall in $R_{m}$, $m = 1, \dots, M$. Given such a partition, for an observation $i$ such that $\mathbf{x}_{i}^{\prime}$ falls in region $R_{m}$, define

equation[equation omitted — 142 chars of source]

such that the predicted probability of $i$ being in class 1 is the proportion of observations in $R_{m}$ belonging to class 1.

Given this rule, ideally we would choose a set of regions in order to minimize our loss function, however this is in general computationally infeasible. Instead, a greedy algorithm is used that in each step chooses a predictor and splits the feature space in two at some value of that predictor. At the beginning of the algorithm, suppose we choose predictor $j$ and some threshold value for that predictor, $s$. These values can be used to split the feature space into two regions:

align[align omitted — 163 chars of source]

Within these regions, observations are classified according to (ref). So we may choose predictor $j$ and splitting value $s$ to minimize the loss function. The algorithm can then be repeated at each step within each newly created region.

Since each region is determined by a series of sequential splits, we can represent the classifier by a tree diagram in which each node is a predictor and the two branches emanating from the node represent which side of the split the value of that predictor falls on. How large the tree grows is a tuning parameter that can be cross-validated.

(ref) illustrates an example of a classification tree.

Random Forest

The predictive performance of a decision tree can be improved at the expense of interpretability by forming an ensemble of trees. Bootstrap aggregated (bagged) decision trees can be formed by training separate decision trees on bootstrap samples of the data Breiman1996. On different realizations of the data from each bootstrap sample, different decision rules can be learned by each tree. Averaging predictions across these trees serves to decrease prediction bias.

However, if a few of the predictors are particularly strong, the decision trees will have a tendency to use these predictors in the early splits on most of the bootstrap samples, leading to small variation in the rules learned by the trees. To combat this, \shortciteA{Breiman2001} suggested only allowing a random subset of predictors to be used on each bootstrap sample to promote variation in the rules learned by the trees.

In addition to the size of each tree being a tuning parameter we now also need to tune the bootstrap sample size, the number of predictors to randomly choose on each bootstrap, and the number of trees in the ensemble. Standard convention is to use bootstrap samples of same size as the data, $n$. Default values for the number of predictors to randomly sample for each tree are $\sqrt{p}$ for classification and $p/3$ for regression Friedman2001. Since random forests do not overfit as the number of trees increases, the number of trees can be as large as the modeler desires.

We train a random forest with out-of-the-box parameter choices for the bootstrap sample size and number of predictors selected in each sample. We train an ensemble with 500 trees, which keeps the compute time manageable given the size of our data set.

Neural Network

Neural networks are an especially popular machine learning tool because they can learn sophisticated non-linear functions when both sample sizes and number of predictors are large. Contrast this with decision trees, which learn relatively unsophisticated non-linear functions due to partitioning the feature space into rectangles, rather than more complicated shapes; and support vector machines or k-nearest neighbors, which become computationally infeasible for large $n$ and $p$. Neural networks have been used to forecast time series, translate text, and learn to play games. They are particularly useful in `deep learning' for computer vision, where they are used classify and label objects in images, automatically generate captions for images, and generate photorealistic images from random noise.

figure[figure omitted — 909 chars of source]
figure[figure omitted — 1,657 chars of source]

Neural networks attempt to imitate the biological function of neurons and axons in a brain. The basic building block of a neural network is a neuron (or hidden unit). Neurons accept a vector of input values, perform a linear transformation on those values, apply an activation function to the output of that transformation, and output the result of the activation function ((ref)). The coefficients (or weights) and intercept term (or bias) in this linear transformation are parameters of the neural network that must be estimated.

The application of an activation function is what grants the neural network non-linearity. The activation function attempts to emulate an action potential that decides whether or not a neuron in the brain `fires'; i.e., passes a large output value on to other neurons. Popular choices for the activation function are the sigmoid function, and the rectilinear function.

align[align omitted — 97 chars of source]

Neurons with a rectilinear activation function are called rectilinear units (ReLU). The neural network is built by connecting a set of neurons together such that the outputs of some neurons are the inputs of others.

In this article we use a popular neural network architecture called a multilayer perceptron (MLP). An MLP structures the neural network as an input layer, a set of hidden layers, and an output layer. The input layer does not contain any neurons and simply feeds the vector of predictors, $\mathbf{x}_{i}^{\prime}$, into each neuron in the first hidden layer. Hidden layers contain a set of neurons that are connected to the previous and subsequent layers, but share no connections within the same layer. There may be many hidden layers, or only one (as in the case of the Single Layer Perceptron), and each hidden layer may have a different number of neurons. The output layer contains a set of neurons whose output corresponds to the output of the function the neural network is estimating and tends to have a different activation function than hidden layers. For regression problems, the output layer is typically a single neuron with a linear activation function; and for a $k$-class classification problem, the output layer contains $k$ neurons, corresponding to the predicted probabilities than an observation is in class $k$, with a softmax activation function to normalize those probabilites such that they sum to one.

(ref) illustrates the MLP structure used here. We use 4 hidden layers with 256, 128, 64, and 32 neurons respectively. Since we are trying to classify a binary variable, the output layer of the neural net contains two units and a softmax activation function. The predicted probability that an observation belongs to the positive class, $\hat{y}_{i}$, corresponds to the output of one of these neurons.

The neural network's weights are estimated by minimizing the loss function using the backpropogation algorithm Rumelhart1985. Formerly, common practice to avoid overfitting was to stop the training of the neural network early according to error on a validation set. However, with the invention of the dropout layer Srivastava2014, overfitting can be avoided while allowing the network to be trained to optimality. Dropout layers accomplish this by randomly preventing a proportion of neurons in each layer in each training batch from propogating their output forward into the next layer. This is somewhat analogous to random forests, where a random subset of predictors are used to build each decision tree. We regularize the network using dropout layers after each hidden layer, with dropout rate set to 20%. The network is trained for 100 epochs.

Model Evaluation

Binary classification models are often evaluated using accuracy (or balanced accuracy) on a test set. Accuracy is calculated as the proportion of correct predictions, where an observation is predicted to be in class 1 if $\hat{y}_{i} \geq 0.5$ and class 0 otherwise. However, in some settings we may wish to err on the side of a classifier being particularly sensitive or specific. For example, in scientific literature low sensitivity is tolerated in order to have high specificity, e.g., through the $p < 0.05$ convention for p-values, in an attempt to avoid false discoveries. Whereas in medical testing high sensitivity may be preferable to avoid failing to detect a serious medical condition. Similarly, in the early detection of academic risk, one might wish to favor sensitivity or specificity in a classifier.

A range of sensitivity/specificity combinations can be achieved for a classifier by classifying an observation into class 1 if $\hat{y}_{i} \geq z$, for some $z \in [0, 1]$. The frontier of sensitivity/specificity combinations that a classifier can achieve by varying $z$ is called the Receiver Operating Characteristic (ROC) curve. In order to remain agnostic with respect to a desirable sensitivity or specificity, we measure model performance using area under the ROC curve (AUC) evaluated on a test set.

Let $C_{1} = \{ i \, | \, y_{i} = 1\}$ be the set of observations not achieving minimum standards and $C_{2} = \{j \, | \, y_{i} = 0\}$ be the set meeting minimum standards, with $|C_{1}| = N_{X}$ and $|C_{2}| = N_{Y} = n - N_{X}$. For a given classifier, let $\{X_{i}\}_{i=1}^{N_{X}}$ be the set of predicted probabilities $X_{i} = \hat{\mathbb{P}}(y_{i} = 1)$ for each $i \in C_{1}$ and let $\{Y_{j}\}_{j=1}^{N_{Y}}$ be defined analogously for $j \in C_{2}$.

Given some threshold $z \in [0, 1]$, the classifier predicts that an observation $i \in C_{1}$ belongs to the positive class if $X_{i} \geq z$, and analogously for $j \in C_{2}$. The classifier makes a correct prediction when $X_{i} \geq z$ or $Y_{i} < z$.

The sensitivity and specificity of the classifier are defined as follows.

align[align omitted — 197 chars of source]

The set of sensitivity/specificity pairs for $z\in[0, 1]$ defines the ROC curve. (ref) illustrates. The area under the ROC curve (AUC) neatly summarizes the set of achieveable sensitivity/specificity combinations. An AUC of 1 is most desirable, corresponding to both a sensitivity and specificity of 1. One classifier with larger AUC than another is able to achieve a greater sensitivity for a given specificity, and vice versa.

figure[figure omitted — 335 chars of source]

\shortciteA{BAMBER1975} shows that when the area under the ROC curve is calculated with the trapezoidal rule, the estimator for AUC is equal to

equation[equation omitted — 115 chars of source]

where

equation[equation omitted — 133 chars of source]

In the case where $X$ and $Y$ are continuous, such that $\mathbb{P}(X = Y) = 0$, the AUC is equivalent to the Mann-Whitney U statistic, which is asymptotically normal with known variance. \shortciteA{BAMBER1975} uses a result from \shortciteA{noether1967} to estimate the variance of AUC when $X$ and $Y$ are not continuous. Using the formulation from \shortciteA{DeLong1988}, define

align[align omitted — 316 chars of source]

The variance for the AUC estimate can then be expressed as

equation[equation omitted — 138 chars of source]

This variance can be estimated by replacing population quantities in (ref) with their sample analogues and confidence intervals for the AUC estimate can then be constructed according to asymptotic normality.

\newcolumntype{L}[1]{>{\let\newline\\\arraybackslash}m{#1}} \newcolumntype{C}[1]{>{\let\newline\\\arraybackslash}m{#1}}

table*[table* omitted — 1,289 chars of source]

Results

ROC curves for each classifier are displayed in (ref). (ref) contains performance metrics for each model. We see that no ML model is able to outperform logistic regression in terms of AUC. The logistic classifier has the highest or equal highest AUC in all cases and in many cases the decision tree, random forest, and neural network perform statistically significantly worse. The elastic net matched the performance of the logistic, indicating that the penalty used to exploit the bias-variance trade-off did not improve predictive ability. Further, the decision tree and random forests significantly underperformed the logistic regression and elastic net.

The reason for ML not delivering substantial improvement in the academic performance setting possibly resides in the nature of common predictors used being categorical. Part of ML's power comes from its ability to tractably model non-linearities in the data. A lack of continuous predictors with substantial non-linearities being present in education data may explain part of this result. However, ML can still give substantial improvement even when most predictors are categorical. Decision trees (and hence random forests) and neural nets can naturally model complex interaction effects between categorical predictors. These results suggest that such interaction effects may not be present in the education setting either.

Since the grade 5 and above data contains past performance predictors, these models perform significantly better than their grade 3 counterparts. The logistic AUCs for the Grade 5+ predictions are 0.839 and 0.833 for literacy and numeracy respectively, compared to only 0.722 and 0.707 for the Grade 3 predictions. While this is to be expected, since the grade 5+ data set has an additional strong predictor, it is desirable to detect underachievement as early as possible due to its persistence.

While the tree-based methods perform significantly worse than logistic regression, they give highly interpretable decision rules. These rules may be useful in terms of interpretation, as they provide simple rules or heuristics for detecting students at risk of poor performance. The potential advantage of such simplicity is that they could easily be learned and applied with minimal effort by school staff. (ref) displays.

When previous achievement on NAPLAN is available (i.e. for Grade 5+), the decision trees learn a very simple and intuitive decision rule. If a student was below standard on either literacy or numeracy in the previous testing period, they are classified as at risk in both subjects in the next period. This is of value, as it indicates that a school faced with a student that has performed below standard in numeracy in one period should not shift all of their resources towards improving only their math ability. In the United States, \shortciteA{Reback2008} highlighted that Texan schools reallocated resources towards students based on their performance on previous tests. We provide evidence that this may not be a useful approach, and instead a focus on overall improvement in all subject areas should be emphasised.

The decision trees for grade 3, for which achievement in the previous period is not available, provide some insight into the predictors of poor achievement in Australia. Categorising students as at risk or not depends completely on the education levels and occupation of their parents. For both literacy and numeracy, the first split is made regarding mother's higher education. If a student's mother holds a Bachelor's degree, a student will be predicted to meet minimum standards. However, if the mother does not hold a degree, their performance depends on the education of their father. If their father also does not have higher education qualifications, the student will be predicted not to meet minimum standards. If their father has either a Bachelor's degree, a diploma or an advanced diploma, the student will be predicted to meet minimum standards. In general terms, if at least one of the student's parents has a Bachelor's degree, or their father has a diploma or an advanced diploma, a student is not at risk of poor performance on NAPLAN, in reading or numeracy. This finding matches much of the literature, which finds that family background strongly determines educational outcomes Cobb2012,Ford2013. In fact, \shortciteA{Nicoletti2013} found that family background explains 44-55% of the variation in test scores for students in England.

The results are the same for literacy and numeracy, save the addition of a decision rule regarding father's occupation for numeracy. If the father is employed, and is not employed in a 'category 4' job, a student will be predicted not at risk of not meeting minimum standards in the numeracy test. Category 4 encompasses machine operators, hospitality staff, assistants, labourers and related workers.

While these decision rules are simple and intuitive, they may be limited in their practical application due to the broad nature of each category. However, these results show the feasibility and potential for these kinds of models for predicting test score outcomes in future should more detailed data become available.

Discussion

The machine learning models we estimated in the paper did not outperform standard logistic regression. Logistic regression was equal best in terms of the point estimate for model performance, so no ML model even managed to achieve a statistically insignificant improvement.

A partial explanation for this could be the relative lack of numerical predictors, for which our data set contained only age.\footnote{We performed the analysis with the two previous NAPLAN achievement variables as numerical scores rather than dummies for `at standard' or `below standard' achievement with no change to the results.} Part of ML's ability to improve predictive power comes from using non-linear estimators of the conditional expectation function, whereas OLS and logistic regression are limited to estimating linear functions (with the log-odds ratio being linear in the case of logistic regression). Having few numerical predictors, or numerical predictors for which the conditional expectation function is close to linear could prevent ML models from performing better.

However, in combining predictors non-linearly, ML models should also be able to learn interaction effects if they are present. If one suspected that interaction effects were relevant in a prediction problem involving OLS or logistic regression, but did not know which interactions predicted well, one could saturate the model with all combinations of two variables. For $p$ predictors, this would involve adding on the order of $p^{2}$ additional predictors. For a large number of predictors or a small sample size, this would become a high-dimensional problem and one would have to resort to machine learning methods anyway, such as regularization or subset selection. Alternatively, one could use a model that can learn interaction effects without having to be given interaction predictors, such as decision trees and neural networks. For our data set, strong interaction effects appear not to be present.

Our result aligns with studies in the educational data mining literature that have benchmarked ML models against logistic regression in the prediction of academic performance Kotsiantis2004, Cortez2008, HUANG2013, Gray2013, although while our study finds no improvement at all from ML, some of these studies find a slight increase in predictive power. These studies have primarily used college level administrative data and have had comparatively small sample sizes. We have shown that this result also holds at the primary and middle school level with a data set of over a million observations. In order for ML to gain traction in this setting, we would likely need data sets richer in numerical predictors, such as scores on formative assessment from classes taken in previous years. However, in the primary and middle school setting, this may not be feasible if letter grades are used more commonly than numeric scores. Additionally, model performance could be improved by collecting predictors that are time varying. Most of the predictors in our model are time invariant, except age, grade, and previous NAPLAN performance.\footnote{Employment status of the parents is time varying, but is only collected at the time of enrolment.}

Finally, any exercise in producing an early warning model for academic risk needs to be mindful of possible data set contamination from unobserved interventions. Presumably the teachers of students in a data set are aware when one of their students has not obtained minimum achievement and may attempt to mobilize school resources to correct for a student that has fallen behind, such as by assigning a Student Support Officer to that student. Whether or not such an intervention has occurred, and how many times an intervention has occurred, is an important variable to include. Otherwise, students who receive the intervention and are able to reach minimum standards could be classified by a model as not at risk when they in fact are at risk absent the intervention.

Conclusion

This paper applied a range of popular machine learning methods to the prediction of academic risk for Australian primary and middle school students as measured by a compulsory national standardized test. No machine learning model was able to significantly outperform standard logistic regression, despite a large data set with many predictors.

The popularity of machine learning is not undeserved. In many areas, the increase in predictive power ML brings is very impressive. However, as we have shown in this article, it is not a panacea. Nor is the `bigness' of a data set.

\titleformat{\section} {\normalfont } {\thesection.}{6pt}

\onecolumn

figure[figure omitted — 926 chars of source]
sidewaysfigure[p] \begin{subfigure}[b]{0.45\textwidth} \caption{Literacy. Grade 3.} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \caption{Numeracy. Grade 3.} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \caption{Literacy. Grade 5+.} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \caption{Numeracy. Grade 5+.} \end{subfigure} \caption{Decison trees. See (ref) for variable descriptions.}
table[table omitted — 4,335 chars of source]

\linespread{1}