Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
68,265 characters · 21 sections · 77 citation commands
Panel semiparametric quantile regression neural network for electricity consumption forecasting
\baselineskip 15pt
\vskip 0.3cm \onehalfspacing
Over the past $40$ years of reform and opening up, China's energy industry has undergone tremendous changes and made remarkable achievements. The total amount of energy production and consumption ranks first in the world. Moreover, the clean energy consumption structure has been greatly optimized, and the total amount of which stands in the front ranks, see CEPPE2018. Energy development has injected a continuous impetus into social and economic development, among which, the electricity, an important secondary energy source, has been widely used and become the material basis of modern civilization since it is easy to clean and control. By the end of $2017$, the installed capacity of electric power in China had reached to $1.778$ billion kilowatts (kW), among the highest number in the world.
However, the electricity resources are not evenly distributed given that China has a vast territory and complex societal, economic and natural conditions, which is reflected in Figure 1, the Annual electricity consumption distributions of $1999$, $2005$, $2010$ and $2017$ in China. In which, the local or regional electricity shortages are highlighted, for example, in 2002, 21 provinces suffered from power shortages; the severest electricity shortage (for about 30 million kW) occurred in 2011 since 2004, which is twice as much as the total power generation of Anhui province, see Yuetal2015. Electricity is also concerned to be a vital driver of economic development. A robust and accurate forecasting model is desirable for policy maker in both developed and developing countries. Consequently, helping to capture the future development trend of economy and the electricity demand, the annual Electricity consumption forecasting (ECF) is imperative. ECF plays a critical role in monitoring and planning the transportation of electric power. In addition, it helps to save energy, reduce pollution, and improve the security and stability of the power system. But ECF is highly affected by many factors such as population, economy growth, industrial structure, income level of residents, climate, national policy, power facilities and so on. It makes ECF a challenging task.
\vskip -1.5cm
Recent years have witnessed the development of various forecasting techniques, which can be divided into four major groups: autoregressive integrated moving average (ARIMA), multivariate linear regression (MLR), grey prediction models and artificial intelligence (AI) learning. As one of the most popular time series models, ARIMA is widely used for ECF. Ediger and Akar Ediger2007 applied ARIMA to predict primary energy demand of fuel in Turkey. Mohamed et al. Mohamedetal2010 considered double seasonal ARIMA model to forecast load demand in Malaysia. Filik et al. Filik2011 proposed a hourly forecast of long-term electric energy demand by ARMA in Turkey. Xu et al. Xuetal2015 introduced a grey model and ARMA (GM-ARMA) method to predict the energy consumption in China. Hussain et al. Hussainetal2016 applied Holt-Winter and ARIMA to forecast electricity consumption in Pakistan. Cabral et al. Cabraletal2017 presented a spatial ARIMA model for ECF in Brazil. Oliveira and Oliveira Oliveira2018 forecasted mid-long term electricity consumption by combining ARIMA with exponential smoothing and bootstrap aggregating technique. MLR is a statistical approach which is widely used for ECF, see for instance, Bianco et al. Biancoetal2013, Meng and Niu MengNiu2011 and Kaytez et al. Kaytezetal2015. This predictor can be carried out easily, but the flexibility and fitting of MLR and ARIMA may not be satisfactory in practice because of the pre-defined constraint. Grey prediction model is also a commonly used tool in electricity consumption prediction, details can be found in Wang et al. Wangetal2018, Ding et al. Dingetal2018, Tang et al. Tangetal2019. As far as we know, AI methods are the highly popular prediction tools, which is able to perform well in dealing with modeling complexity and nonlinearity. Different AI methods have been proposed for electricity consumption/demand/load prediction, for example, Artificial Neural Networks (ANNs) (Azadeh et al.Azadehetal2008a, Azadeh et al. Azadehetal2008b, Kandananond Kandananond2011, He et al. Heetal2019), Support Vector Machine (SVM) (Pai and Hong PaiHong2005, Fan and Chen FanChen2006, Hong Hong2009), Support vector regression (SVR) (Elattar et al. Elattaretal2010, Yang et al. Yangetal2019) and Genetic programming (GP) (Mostafavi et al. Mostafavietal2013), among others. However, these techniques work well only in time series or cross-sectional data, not the panel data.
In this paper, we propose a novel approach called Panel Semiparametric Quantile Regression Neural Network (PSQRNN) to analyze and forecast the provincial electricity consumption in China by ANN and semiparametric quantile regression (SQR). Quantile regression (QR) of Koenker and Bassett KoenkerBassett1978 explores the relationship among variables comprehensively, including median regression as a special case. Since quantile estimator can quantify the entire conditional distribution of the response variable conditional on covariates and give a global assessment of the covariate effects at different quantile of the response (Koenker Koenker2005). It is particularly useful when the conditional distribution is heterogeneous and asymmetric, heavy-tailed, or truncated in Shimetal2012. However, most research of QR (KoenkerKoenker2005) with multiple predictors have relied on linear or parametric nonlinear models, see Cannon Cannon2011. To take the advantage of QR and ANN, Taylor Taylor2000 introduced a more flexible model: Quantile Regression Neural Network (QRNN), which can implement a nonlinear QR and grasp nonlinear relationship between the response and covariates without specifying a precise functional form. Related work includes Xu et al. Xuetal2016, He and Li HeLi2018, He et al. Heetal2019. Further, Xu et al. Xuetal2017 used a composite quantile regression neural network (CQRNN) model to explore the potential nonlinear relationship among variables. Cannon Cannon2018 considered non-crossing nonlinear regression quantiles by monotone CQRNN, called MCQRNN; Jiang et al. Jiangetal2017 developed an expectile regression neural network by adding ANN to expectile regression, which is similar to QRNN. Since QRNN model is a purely nonlinear QR, the flexibility and good forecasting performance is guaranteed. However, QRNN is considered as a black-box model due to the weakness of interpretability.
Recently, semiparametric regression model has become popular because it keeps the flexibility of nonparametric models and maintains the interpretability of parametric models simultaneously (Kai et al. KaiLiZou2011). The application of semiparametric model is dated back to Engle et al. Engleetal1986, Fan and Hyndman FanHyndman2012, Weron and Misiorek WeronMisiorek2008, Shao et al. Shaoetal2014, Goude et al. Goudeetal2014, Shao et al. Shaoetal2015. In the background of semiparametric quantile regression (SQR), Lebotsa et al. Lebotsaetal2018 considered a short term electricity demand forecasting using partially linear additive QR.
We focus on the heterogeneity of the provincial electricity consumption based on a cross-province study. In conditional mean panel data (linear) models, taking a difference is commonly used to eliminate the individual effect, but it is invalid in the background of QR, particularly for linear QR model. But such study is rare in literatures (Cai et al. Caietal2018). Koenker Koenker2004 first introduced a panel quantile regression (PQR) which treats the individual fixed effects as a pure location shift parameters common to all conditional quantiles. Later, Lamarche Lamarche2010 contributed on studied theoretical properties of PQR, whereas Calvao Galvao2011 extended the quantile regression to a dynamic panel data model with fixed effects. The application of PQR includes: Chen and Lei ChenLei2018 revisited the environment-energy-growth nexus by employing a PQR to incorporate the effects of renewable energy consumption and technological innovation; Wang et al. WangZhuetal2018 investigated the effect of democracy, political globalization, and urbanization on PM 2.5 concentrations with G20 countries based on evidence from PQR; Wang et al. WangZengLiu2019 applied a PQR and a balanced city panel data in China to examine the multiple impacts of technological progress on CO$_2$ emission. However, such PQR models are linear panel quantile regression. To address the issue of nonlinearities and heterogeneity simultaneously, Cai et al. Caietal2018 proposed a semiparametric quantile panel data model with correlated random effects, in which some of the coefficients are allowed to depend on smooth economic variables while the other coefficients are constant, to estimate the growth effect of foreign direct investment, where the nonparametric component is used to model nonlinearity, but it is not flexible enough compared to neural network. Moreover, nonparametric nonlinear model suffers from model misspecification, which leads to an inaccurate forecast. Motivated from which, we propose the PSQRNN to forecast the provincial panel electricity consumption in China.
The contributions of our paper are: (1) PSQRNN is proposed for the regional electricity consumption forecasting, which combines panel data, semiparametric model and composite QR with QRNN. (2) PSQRNN is a new framework by adding parametric structure and unobserved regional heterogeneity to QRNN, which can explore potential linear and nonlinear relationship among the variables and interpret the unobserved cross-sectional heterogeneity simultaneously, and maintain a better interpretability of parametric models. (3) A new estimator is derived to train the PSQRNN model by assembling the penalized quantile regression with LASSO, ridge regression and backpropagation algorithm.
The rest of the paper is organized as follows. In Section 2, we introduce the PSQRNN model and address the issues of corresponding estimation and selection. Section 3 gives the comparison of PSQRNN with three competitive intelligent methods including BP neural network (BP), SVM and QRNN, and presents the provincial electricity consumption forecasting in China under three scenarios via training, evaluating and forecasting. Conclusions are drawn in Section 4.
To explore the relationship between a response $y$ and predictors $X=(x_1,\cdots,x_p)^T$ for a time series and cross sectional data, the linear quantile regression (LQR) model can be written as
where $Q_\tau(y_i|\vc X_i)$ is the $\tau$th condition quantile of $y_i$, $\{(\vc X_i,y_i), i=1,\cdots,n\}$ is the observed data, $\vc X_i=(x_{i1},\cdots, x_{ip})$, $n$ is sample size, $\beta_\tau$ is regression coefficients, and $\tau\in(0,1)$, which follows the settings of Koenker Koenker2005. This model only considers the linear relationship between the response variable and the predictors under different quantiles, which is inappropriate to investigate the nonlinear relationship in practice. Thus, Cannon Cannon2011 proposed a QRNN model in term of one hidden layer,
with
where $w_j^{(h)}$ and $b^{(h)}$ are the hidden-layer weights and bias, $w_j^{(o)}$ and $b^{(o)}$ are the output-layer weights and bias, and $J$ is the number of node of the hidden layer, $a(\cdot)$ is the active function and $f(\cdot)$ is the output-layer transfer function. One can refer to the schematic diagram in Figure 2 (a), which shows a QRNN model with two hidden layers.
In this paper, we study a panel data of provincial electricity consumption. Let a scalar dependent variable $Y_{it}$ be the observation of $i$th individual at time $t$ for $i=1,\cdots N$ and $t=1,\cdots,T$. The panel semiparametric quantile regression model (PSQR) is:
where $Q_\tau(Y_{it}|\vc X_{it},\vc Z_{it},U_{it},\alpha_i)$ is the $\tau$th quantile of $Y_{it}$ given $\vc Z_{it}, \vc X_{it}, U_{it}$ and $\alpha_i$, $\vc Z_{it}$ and $\vc X_{it}$ are $p\times1$ times $q\times 1$ predictors, respectively. $\vc\beta_\tau$ is a constant regression coefficient, $\vc\gamma_\tau(U_{it})$ denotes a regression functional coefficient of $U_{it}$, which is also an observable scale predictor, and $\alpha_i$ is an individual effect. See Cai et al. Caietal2018.
\vskip -1.5cm
The model PSQR ((ref)) is a traditional statistical model, which combined LQR for panel data (Koenker Koenker2004) and nonparametric regression model. To avoid the “curse of dimension" in pure nonparametric model, a varying coefficient model is embedded into the PSQR. The varying coefficient $\vc X_{it}^T\vc \gamma_\tau(U_{it})$ in model ((ref)) concerns the nonlinear relationship of the response $Y$ and the predictors $(\vc X, U)$. The nonparametric nonlinear form and predictions are forcibly divided into $\vc X$ and $U$, with a high risk of model misspecification. In order to avoid an inaccurate forecasting, we propose a new panel semiparametric quantile regression model based on an artificial neural network.
Motivated by LQR and QRNN, we develop a new learning method to predict provincial electricity consumption in China, which combines PSQR and ANN. We call it PSQRNN, which is designed under a multi-layer perceptron (MLP) framework. To begin with, a general dependent structure of PSQRNN is:
where $ANN(\cdot)$ is an arbitrary nonlinear function, whose structure is simulated by utilizing ANN, and $\vc\theta_\tau$ is a parameter of weights and biases in ANN. Let the input and output layer is the $0$th layer and the $L+1$th layer, $\vc g_{it}^{(0)}=\vc X_{it}$, $\vc g_{it}^{(l)}=(g_{1,it}^{(l)},g_{2,it}^{(l)},\cdots,g_{n_l,it}^{(l)})^T$ is a vector of $n_l$ nodes at the $l$th hidden layer for $l=1,\cdots,L$, i.e., $\vc g_{it}^{(l)}\in R^{n_l\times1}$, while the weighted matrix $\vc W^{(l)}\in R^{n_{l-1}\times n_l}$ and the hidden-layer bias $\vc b^{(l)}\in R^{n_l\times1}$.
First, at the input layer,
At the hidden layer,
where $a^{(l)}(\cdot)$ is an activation function, which controls the network's nonlinearity, maps the real line to its subset, and often adopts forms like signmoid($\cdot$), tanh($\cdot$), ReLU($\cdot$), softplus($\cdot$), etc.
At the output layer, an estimator of the $\tau$th conditional quantile of response is then given by
where $f(\cdot)$ is the output-layer transfer, which is usually an identity function. Here, $ANN(\cdot)={\vc W^{(L+1)}}^T\vc g_{it}^{(L)}$. The schematic diagram is depicted in Figure 2 (b), which shows a PSQRNN model with two hidden layers.
In this section, we first introduce a penalized quantile regression model for panel data, then the estimation procedure of PSQRNN is proposed.
In model ((ref)), if the nonlinear $ANN(\cdot)==0$, the PSQRNN becomes a linear quantile panel data model,
which is proposed by Koenker Koenker2004, where $\alpha_i$ is treated as fixed effects with pure location shift effects on the conditional quantiles of the response variable, but the effects of regressions depends on the quantile. While in penalized quantile regression, the quantile loss function is minimized by adding $\ell_1$ penalty on fixed effects. The objective function is:
where $\rho_\tau(u)=u(\tau-I(u<0))$ is the quantile loss function. The weights $w_k$ controls the $K$ quantiles $\{\tau_1,\cdots,\tau_q\}$ when estimating parameters $\alpha_i$, and $\lambda$ is the tuning parameter which reduces the individual effects to improve the accuracy and robustness of the estimation of $\beta$. In terms of which, Lamarche Lamarche2010, Galvao Galvao2011 and Canay Canay2011 studied the penalized quantile regression model for panel data ((ref)) from different perspectives.
Inspired by the estimation of PQRPD in ((ref)) and the implementation of QRNN in Cannon Cannon2011, we define the loss function of PSQRNN as
where $\hat{Q}_{\tau_k}(Y_{it}|\vc X_{it},\vc Z_{it},\alpha_i)$ is given in ((ref)), $\theta=\left\{\alpha_i,W^{(l)},b^{(l)}|i=1,\cdots,N,l=1,\cdots,L\right\}$, $W^{(l)}=(w_{ij}^{(l)})_{n_{l-1}\times n_l}$, $N_L=\sum_{l=1}^Nn_{n-l}n_l$, and $\lambda_1$, $\lambda_2$ are the model penalty parameters, respectively. Here, we apply $\ell_1$ shrinkage to individual effects $\alpha_i$ and conventional Gaussian $\ell_2$ penalties to weights $w_{ij}^{(l)}$. To improve the accuracy of prediction, $\lambda_1$ shrinks the individual effects estimators toward zero and $\lambda_2$ effectively prevents the model from overfitting.
Typically, the loss $\mathcal{L}(\theta)$ with weights and biases of ANN is minimized by gradient descent (GD) and backpropagation algorithm. However, the check functions $\rho_\tau(\cdot)$ and $|\cdot|$ usually are non-differentiable, which is obvious since the derivative is not valid at the origin. Instead, one can replace $\rho_\tau(\cdot)$ and $|\cdot|$ with approximations that are differentiable everywhere. Following Chen Chen2007 for quantile regression and Cannon Cannon2011 for QRNN, the Huber norm, which provides a smooth transition between absolute and squared errors around the origin, is defined as follows:
which is employed to approximate $\rho_\tau(\cdot)$ and $|\cdot|$ by
and $|\alpha_i|=h\varepsilon(\alpha_i)$, where $\varepsilon$ is a pre-deterministic threshold, with $\varepsilon=2^{-i}$ for $i=-8,-9,\cdots,-32$, which is default in the R package qrnn. Hence, the standard GD optimization algorithm is employed to optimize the approximate loss function in the following
This procedure is conducted by using nlm for Newton-type algorithm or optim for Nelder-Mead, and quasi-Newton, conjugate-gradient algorithms in R package. In the entire optimization procedure, $\varepsilon$ in $\mathcal{L}^\varepsilon(\theta)$ begins with a larger starting value, and is updated in each iteration. The algorithm converges until $\varepsilon$ goes to zero.
The PSQRNN model is flexible and useful to reveal the nonlinear predictor-predicted relationship, and contains some latent interactions between predictors. The complexity of ANN in PSQRNN model is determined by $(p,L,n_1,\cdots,n_L)$. A model that is too complex may result in over-fitting, but can be avoided by penalizing a larger weight in the input-hidden layer by adding a quadratic penalty term, which has been considered in our model. $\lambda_2$ in ((ref)) contributes to the terms of weight decay.
Another issue of PSQRNN modeling is to choose $(p,L,n_1,\cdots,n_L)$ and $(\lambda_1,\lambda_2)$, which play an important role in training and prediction. In practice, $p$ is a fixed parameter since predictors in ANN are predetermined. However, to find an optimal combination is unrealistic due to the computational consuming. In practice, $L$ is taken to be $1$ or $2$. As a criterion for model selection, the Bayesian information criterion (BIC) is applied in this paper. For $L=1$, one defines YQ$_{it1}$=$Y_{it}-\hat{Q}_{\tau _{k}}(Y_{it}|{\mbox{\boldmath${X}$}} _{it},{\mbox{\boldmath${Z}$}}_{it},\alpha _{i};n_{1},\lambda _{1},\lambda _{2})$, and
The optimal values of hyperparametrics $(n_1, \lambda_1,\lambda_2)$, $(n_1^*, \lambda_1^*,\lambda_2^*)$ are determined by
For $L=2$, one defines
where YQ$_{it2}$=$Y_{it}-\hat{Q}_{\tau _{k}}(Y_{it}|{\mbox{\boldmath${X}$}} _{it},{\mbox{\boldmath${Z}$}}_{it},\alpha _{i};n_{1},n_{2},\lambda _{1},\lambda _{2})$, and
In general, grid search method can be used to minimize $BIC$, which is simply an exhaustive searching method through a manually specified subset of the hyperparameter space. In the empirical analysis, we adopt two hidden layers, that is, $L=2$. The four-dimensional gird search method for the case of $L=2$ is time-consuming. In reality, $n_1$ and $n_2$ are predesigned, and a two-dimensional grid search method is used to choose an optimal $(\lambda_1^*,\lambda_2^*)$. We find that the PSQRNN model with a small number of nodes fits and predicts well, when $L=2$.
A sketch of training and prediction is provided in the following.
$\bullet$ The selection of $K$, $\tau_k$ and $w_k$. In ((ref)) and ((ref)), fitting multiple values of $\tau$ simultaneously allows one to “borrow strength" from regression quantile and improve the global model performance (Cannon Cannon2018). Typically, $\tau_k=k/(k+1)$, $k=1,\cdots,K$ are equally spaced. $w_k$ are weights that allow regression quantiles for each $\tau_k$ to contribute to the total error. Constant weights $w_k=1/K$ yield a standard composite quantile regression error function in ((ref)) and ((ref)). We set $w_k=1/K$ since there is no prior information. Koenker Koenker2004 pointed out that the choice of $w_k$ and $\tau_k$ is analogous to the choice of discretely weighted $L$-statistics, see for instance, Mosteller Mosteller1946. $K$ is generally selected by experience such as $3$, $5$ and $9$. For a seriously skewed data, $K$ is selected to be $15$ or $20$.
\vskip -1.1cm
$\bullet$ The selection of $a^{(l)}(\cdot)$ and $f(\cdot)$. The activation function of the hidden layer $a^{(l)}(\cdot)$ yields the nonlinearity of the network, which is usually taken to be $\mathrm{sigmoid}(x)=1/(1+e^{-x})$, $\mathrm{tanh}(x)=(e^x-e^{-x})/(e^x+e^{-x})$, $\mathrm{softplus}(x)=\log(1+e^x)$ and $\mathrm{ReLU}(x)=\max(0,x)$. ReLU is the most popular activation function for neural networks, and is successfully applied to deep neural networks. The greatest advantage of ReLU is to alleviate the vanishing gradient problem (Clevert et al. Clevertetal2016). Recently, Leaky ReLU (LReLU), which is a variant of ReLU, has been shown to be superior to ReLU (Maas et al. Maasetal2013). In our model training, the activation function is $a^{(l)}(x)=\mathrm{ELU}(x)$ for classification with $\alpha=1$ for $l=1,\cdots,L$, and the output-layer transfer $f(x)=x$ for regression.
$\bullet$ Avoid local minima and saddle points. $\mathcal{L}^\varepsilon(\theta)$ in ((ref)) is a continuous non-convex loss function. Minimizing a non-convex objective function is a big challenge for science and engineering. Gradient descent or quasi-Newton methods are ubiquitously used to carry out such minimizations. There are many techniques to avoid local minima and saddle points (Dauphin et al. Dauphinetal2014), specially in the deep neural networks. In this paper, we repeat to train the PSQRNN model by generating different weights and bias as initial parameters, then take the minimum point.
The sketch of the model training is depicted in Figure (ref).
\setcounter{equation}{0}
Many works have explored the influencing factors of electricity consumption, which can be roughly divided into economic factors (He et al. Heetal2019) and climatic factors (Fan et al. Fanetal2019). The impact of the two factors regarding to electricity consumption is extensively studied. However, few of them pay attention to the nonlinear case. In this paper, PSQRNN is applied to analyze and forecast China's electricity consumption. We choose economic factors as Gross Regional Product (GDP), Value-added of Secondary Industry (VASI), Total Retail Sales of Consumer Goods (TRSCG), Total Imports and Exports (TIE), and the climatic factors as Annual Average Temperature (AAT), Annual Average Relative Humidity (AARH), Days of Precipitation ($\geq 0.1mm$) per year (DP) and Sunshine Hours (SH). The panel data of China's $30$ provinces during 1999-2018 is collected, in which dataset of EC, GDP, VASI, TRSCG and TIE are from National Bureau of Statistics of China: {\color{blue} http://www.stats.gov.cn/tjsj/ndsj}; dataset of AAT, AARH, DP and SH are from National Meteorological Information Center of China: {\color{blue} http://data.cma.cn}. Among which, the data of economic factors are complete, but the data of climate factors are missing, so we deal with it by imputation via interpolating the mean. The dataset used in PSQRNN model is summarized in Table (ref).
{{1.6em}{
}}
We do some descriptive statistical analysis of the dataset, and find that (i) These panel data are complex with many inherent or latent tends; (ii) Based on Skewness, Kurtosis and p-value of Jarque-Bera test per year, the distributions of EC, GDP, VASI, TRSCG and TIE are skewed (Skewness$>0$), which is more concentrated than normal distribution with longer tails (Kurtosis$>3$). This indicates the non-normality of the unconditional distribution of these factors (p-value$<0.05$). Thus a quantile regression method is desirable to describe the heterogeneity of the factors in electricity consumption. (iii) The distributions of AAT, AARH, DP and SH per year can not reject the normality, but there are great differences among provinces. In addition, the five indicators increase and the variation tends to be larger year by year; TIE deviates much more from the nominal level, which also poses a difficulty for prediction.
In this subsection, the performance of the PSQRNN for ECF is investigated by empirical analysis. Of interest is the comparison of PSQRNN with three competitive intelligent methods including BP neural network (BP), SVM and QRNN. To evaluate the prediction, the provincial EC and the provincial influence factors (GDP, VASI, TRSCG, TIE, AAT, AARH, DP and SH) in 1999-2013 (15 years) are employed to train the four models (PSQRNN, BP, SVM and QRNN), and the provincial influence factors in 2014-2018 (5 years) are treated as testing data to predict the provincial EC in 2014-2018.
In our PSQRNN model, the economic factors are partially linear with EC via the parametric main effect, while economic factors and climate factors are nonlinear with EC via the ANN. Let $\vc Z=(GDP, VASI, TRSCG, TIE)$, $\vc X=(\vc Z, AAT, AARH, DP, SH)$ and $Y=EC$. The PSQRNN model is
The training data is $\{(Z_{it},X_{it})\rightarrow Y_{it}, i=1,\cdots,30, t=1,\cdots,15\}$. Here and after, “$i$" indicates the province, and “$t$" denotes the year. “$i=1$" is Beijing, $\cdots$, “$i=30$" is Xinjiang, “$t=1$" is 1999, “$t=2$" is 2000, $\cdots$, “$t=15$" is 2013, and so on. Then we use the testing data $\{(Z_{it},X_{it}), i=1,\cdots,30, t=16,\cdots,20\}$ to predict $\{Y_{it}, i=1, \cdots, 30, t=16, \cdots, 20\}$ by the trained PSQRNN model, to obtain $\{\hat{Y}_{it}, i=1, \cdots, 30, t=16, \cdots, 20\}$. We choose $\tau_k=0.01+0.02k$ for $k=0,\cdots,49$ so that $\tau_k\in [0.01,0.99]$, and the number of hidden nodes in the first and second hidden layers $HL=(10,5)$. In order to evaluate the performance of PSQRNN extensively, other number of nodes are utilized in the hidden layers, for which, one can refer to subsection 3.3.
For BP and SVM methods, the input is $\vc X$ and output is $EC$. The procedure of BP is carried out by using the \verb"neuralnet" function in \verb"R" package \verb"neuralnet" (Fritsch et al. Fritschetal2019). The setting scenarios are: the backpropagation algorithm, sum of squared errors calculation, tangent hyperbolicus-type activation function, and the default setting. We use a two-hidden-layer network with $HL=(5,5)$, not $HL=(10,5)$ because it was overfitting. The SVM method is utilized by the \verb"svm" function in \verb"R" package \verb"e1071" (Meyer et al. Meyeretal2019), where the default settings are adopted, for example, the kernel used in training and predicting is radial basis $\exp(-\gamma|u-v|^2)$ with the default $\gamma=3$, and so on. For QRNN, it is modeled as
without semiparametric and individual effect terms in contrast to the PSQRNN model ((ref)), where $i=1,\cdots,30$ and $t=1,\cdots,15$. We use the \verb"qrnn2" in \verb"R" package \verb"QRNN" (Cannon Canay2011), which is developed to fit and predict from QRNN models with two hidden layers network. We also use the number of hidden nodes $HL=(5,5)$, not $HL=(10,5)$ because of overfitting. The quantile probability is $\tau=0.5$, which is corresponding to the least absolute deviation regression, since it is robust to the heavy-tailed data or outliers.
From two perspectives of province and year, the mean absolute percentage error (MAPE) and relative root mean square error (RRMSE) are utilized to be the measurement of forecasting performance (out of sample), for the 30 provinces:
and over the 2014-2018 years:
and
In addition, we also calculate the means and standard deviations (Std Dev) of MAPE$_i$ and RRMSE$_i$ for $i=1,\cdots,30$. These results are listed in Tables 2 and 3, respectively.
{{1.6em}{
}}
{{1.6em}{
}}
Generally, the smaller the MAPEs and RRMSEs are, the better the method performs. Bold values indicate the best performance. From Tables (ref) and (ref), the proposed PSQRNN method performs best or being close to the best one in provinces and years based on MAPEs and RRMSEs. According to means and standard deviations of MAPEs and RRMSEs in Table (ref), it can see that our PSQRNN is more robust than the BP, SVM and QRNN. From Total$\cdot$MAPEs and Total$\cdot$RRMSEs in Table (ref), for example, 0.2621(BP)$>$0.2559(QRNN)$>$0.2345 (SVM)$>$0.1363(PSQRNN) based on Total$\cdot$MAPEs, it is also clear to find that state of the art approaches such as BP, SVM and QRNN are generally inferior ability than our PSQRNN in forecasting for this panel data. Therefore, one finds that the PSQRNN has an excellent prediction accuracy and robustness, while the BP, SVM and QRNN are not applicable since they did not take into account the inherent structure of the panel data.
Our framework of ECF based on the influencing factors via PSQRNN model is constructed in Fig. 2(c). Recall that $\vc Z=(GDP, VASI, TRSCG, TIE)$, $\vc X=(\vc Z, AAT, AARH, DP, SH)$ and $Y=EC$ in our PSQRNN model ((ref)). In the following subsection, we will train the PSQRNN model and forecast the provincial EC in China under three different scenarios. The predictive performance of the PSQRNN is evaluated again based on the first two scenarios. The efforts of the first two scenarios provide a sufficient guarantee for the third scenario to be accurately applied to prediction.
In order to illustrate the importance of our five-year ECF, three scenarios are considered for the five-year forecasting of EC, which is depicted in Fig. (ref). The scenarios are:
Scenario 1: The provincial EC and the provincial influence factors (GDP, VASI, TRSCG, TIE, AAT, AARH, DP and SH) in 1999-2013 (15 years) are employed to train the PSQRNN model ((ref)), and the provincial influence factors in 2014-2018 (5 years) are treated as testing data. The scenario has been used to comparison of prediction method in Subsection 3.2. The MAPEs and RRMSEs are applied to assess the learning/predictive performance on in-sample/out-of-sample of the PSQRNN.
Scenario 2: The provincial EC in 2004-2013 (10 years) and the provincial influence factors in 1999-2008 (10 years) are employed to train the PSQRNN model ((ref)), and the provincial influence factors in 2009-2013 (5 years) are testing data to predict the provincial EC of 2014-2018 (5 years). The learning/predictive precision of the PSQRNN are also evaluated by MAPEs and RRMSEs on in-sample/out-of-sample.
Scenario 3: The provincial EC in 2009-2018 (10 year) and the provincial influence factors in 2004-2013 (10 years) are employed to train the PSQRNN model ((ref)), and the provincial influence factors in 2014-2018 are testing data to forecast EC of the next five years 2019-2023.
In Scenario 1, the PSQRNN model still is ((ref)). The training data is $\{(Z_{it},X_{it})\rightarrow Y_{it}, i=1,\cdots,30, t=1,\cdots,15\}$ and the testing data is $\{(Z_{it},X_{it}), i=1,\cdots,30, t=16,\cdots,20\}$. To predict $\{Y_{it}, i=1, \cdots, 30, t=16, \cdots, 20\}$ by the trained PSQRNN model, and then get $\{\hat{Y}_{it}, i=1, \cdots, 30, t=16, \cdots, 20\}$.
The setting of Scenario 1 is to evaluate the performance of the estimation and prediction. The numbers of hidden modes in the first and second layers are $HL=(5,5)$ and $HL=(10,5)$, respectively. And $\tau_k=0.01+0.02k$ for $k=0,\cdots,49$ so that $\tau_k\in [0.01,0.99]$. We compute $\mathrm{MAPE}_t$ and $\mathrm{RRMSE}_t$ for $t=1,\cdots,20$ as the evaluation criteria over 1999-2013 (in-sample) and 2014-2018 (out-of-sample). The results are listed in Table (ref). From which, the MAPEs and RRMSEs of out-of-sample are small. It shows that the PSQRNN model is well trained and the estimates are accurate. In addition, the MAPEs and RRMSEs of in-sample are $10$-$20\%$, which indicates a good forecasting performance (Dingetal2018). Moreover, the estimation and forecasting performance tends to be better as the network nodes increasing properly.
{{1.6em}{
}}
In Scenario 2, the parametric $\vc Z=(GDP, VASI, TRSCG, TIE)$, the nonparametric $\vc X=(\vc Z, EC, AAT, AARH, DP, SH)$ and the corresponding output $Y=EC$. Based on ARIMA, PSQRNN is implemented in the following
The training data is $\{(\vc Z_{it},\vc X_{it})\rightarrow Y_{i,t+5}, i=1,\cdots,30, t=1,\cdots,10\}$. Then we use the testing data $\{(\vc Z_{it}, \vc X_{it}), i=1,\cdots,30, t=11,\cdots,15\}$ to predict $\{Y_{it}, i=1, \cdots, 30, t=16, \cdots, 20\}$ by the trained PSQRNN model to obtain $\{\hat{Y}_{it}, i=1, \cdots, 30, t=16, \cdots, 20\}$. In order to forecast better, the historical EC data is used as the nonlinear predictive variable.
In PSQRNN training, we follow the same settings in Scenario 1, and calculate two evaluation criteria MAPE and RRMSE as well, which is reported in Table 5. It is obvious that the proposed model ((ref)) performs well for ECF.
{{1.6em}{
}}
In Scenario 3, our goal is to predict the provincial electricity consumption of China during 2019-2023. Since there is no data of GDP, VASI, TRSCG, TIE, AAT, AARH, DP and SH for 2019-2023, model ((ref)) is infeasible. A historical data is utilized to predict the future electricity consumption and the candidate Model ((ref)) is designed as follows: the training data is $\{(\vc Z_{it},\vc X_{it})\rightarrow Y_{i,t+5}, i=1,\cdots,30, t=6,\cdots,15\}$, and the testing data is $\{(\vc Z_{it}, \vc X_{it}), i=1,\cdots,30, t=16,\cdots,20\}$ to predict $\{Y_{it}, i=1, \cdots, 30, t=21, \cdots, 25\}$ (ECF of 2019-2023). The setups are similar to Scenario 3, but there is only two hidden layers, i.e., $HL=(15,5)$ with $(\lambda_1, \lambda_2)=(0.005,0.01)$ being applied to the ANN.
First of all, the performance of forecasting is evaluated. The MAPE$_t$s and RRMSE$_t$s are calculated based on $Y_{it}$ and $\hat{Y}_{it}$, $i=1,\cdots,30$ and $t=16,\cdots,20$ (Years 2014-2018), which are listed Table (ref) with an excellent performance (Dingetal2018) in terms of in-sample forecasting. Figure (ref) is a scatter plot of the ECs and the corresponding predicted provincial values of 2014-2015 and 2017-2018, which reveals an accurate prediction. Second, the estimators of the coefficients $\beta$ in model ((ref)) are $\hat{\beta}=(0.1119, 0.1399, 0.2928, 0.2421)$. It shows that GDP, VASI, TRSCG and TIE have a positive linear effect on EC. Finally, the ECs of 2019-2023 are predicted in Figure (ref) based on the trained PSQRNN model and the historical data. The curves in Figure (ref) are potential trends of electricity consumption demand of China's provinces in the future. The trends vary from province to province, for example, the electricity consumption demand highly in Guangdong, Jiangsu, Shandong and Zhejiang, with a relative steep growth, while the electricity demands in provinces like Sichuan, Hubei, Hunan and Anhui increase steadily and that of provinces such as Hainan, Qinghai, Ningxia, Helongjiang and Gansu grows slowly. These findings are helpful for government decision-making.
\vskip 2cm {{1.6em}{
}}
Electricity is an important ingredient for the economic growth and development of countries. Annual electricity consumption forecasting plays a vital role in the power investment planning and energy development. However, it is particularly challenging to predict the annual electricity consumption for provinces in China, given that China has a vast territory and the electricity consumption shows great heterogeneity in provinces with different economic, social and climatic conditions. Motivated by which, we propose a new model--PSQRNN. It combines panel data, semiparametric model and composite QR with QRNN, which keeps the flexibility of nonparametric models and maintains the interpretability of parametric models simultaneously. In order to train the PSQRNN, a penalized composite quantile regression with LASSO, ridge regression and backpropagation algorithm is considered. In addition, a differentiable approximation of the quantile loss function is adopted so that the quasi-Newton optimization can be conducted to estimate the model parameters.
The prediction accuracy is evaluated by an empirical study of 30 provinces' dataset from 1999 to 2018 in China under three different scenarios, based on economic and climate factors. Compared to the QRNN model, the PSQRNN model is more robust and performs better for electricity consumption forecasting. Finally, the PSQRNN model is used to predict the electricity consumptions of provincial data over 2019-2023. We predict that China's electricity consumption will reach 9082.1 billion KWH by 2023, which is 7.36 and 1.21 times that of 1999 and 2018, respectively. We find that the trends of electricity consumption demand in the future varies from province to province, which is helpful for government decision-making.
There are some interesting future works: (1) For large dimensional panel data, some feature selection techniques, such as principal component analysis, factor analysis and Lasso, can be applied to our proposed PSQRNN. They can screen out more important factors from the social, economic and climatic factors associated with electricity consumption. Therefore, PSQRNN based on feature selection is worth studying for panel electricity consumption forecasting. (2) Build hybrid methods based on Machine Learning, for example, Boosting PSQRNN, PSQRNN with random forest, PSQRNN with GANs, etc. The above methods are supposed to improve the accuracy of prediction.
This work is supported by National Social Science Fund of China (No. 19BTJ034)