Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,082 characters · 22 sections · 71 citation commands
Predicting popularity of EV charging infrastructure from GIS data
Energy consumption has increased outstandingly in the last years and it will continue to be a significant global challenge. The largest portion of the total energy consumption is in the transportation sector where high economic and population growth causes a rapid increase in energy demand with excessive CO\textsubscript{2} emissions and energy crisis EIA2013. To mitigate the impact of the emissions of greenhouse gases and to increase energy security, automotive powertrain electrification in transportation sector could play an important role Tali_2013, wilbanks2010. This is confirmed by the fact that several countries have announced an ambition to stop selling vehicles fueled by diesel or petrol. For example, the UK set as the target year 2040. The EU has taken a decisive step forward in implementing the EU’s commitments under the Paris Agreement for a binding domestic CO\textsubscript{2} reduction of at least 40% until 2030 JRC_web. Electric vehicles (EVs), as a key element of clean and green travel mode, are spreading all over the world rapidly. For instance, an ambitious target of having over 20 million EVs on the roads by 2020 has been set by the U.S. Department of Energy EIA2013. However, the adoption of EVs on a large scale is supposed to bring both challenges and opportunities from technical and economic points of view ajanovic2016dissemination,Flammini2017.
In order to boost EV popularity, many challenges need to be addressed. The chicken-egg problem in the form of charging infrastructure versus EV adoption has been recognized as an important challenge restraining the growing EV ecosystem van2013data. Currently, insufficient charging infrastructure is a significant factor that prevents larger penetration of EVs Coffman2017. The drivers are hesitating to buy an EV if there is not sufficient charging infrastructure, and similarly charging infrastructure operators do not invest while the number of EVs is low and not profitable. In recent years, demand-driven and strategic rollouts (i.e. strategically covering the geographic space by chargers) have been applied Helmus2018. The opening of a new public charging infrastructure involves the estimation of the visitation patterns to ensure as high as possible utilization to justify the allocated resources. Prior to opening a new public charging infrastructure, the expected utilization should be estimated in order to justify the allocated resources. Hence, one way how to support decision making in this area is to develop data-driven approaches with a predictive power.
Different methods, such as mathematical programming, computer-based simulation and statistical learning, have been proposed to deploy EV charging infrastructure (EVCI). Mathematical programming models for the optimal location of EVCI have been proposed in several papers. The methodologies are based on the minimization of objective functions, addressing infrastructure development costs Liu2012, Sadeghi2014, social costs ren2019location, driving range and driver habits Gonzalez2014,Dong2014,He2018, traffic flow data Efthymiou2017, unmet demand Chen2013, and quality of service Davidov2017. In GE2012 was addressed the problem considering road network and distribution system network capacity constraints. The work of Asamer et al. Asamer2016 provides a formulation of the optimal location problem for a specific category of vehicles (taxi service). A comprehensive review of different optimization techniques for EVCI can be found Rahman2016.
In contrast to the large number of studies that applied mathematical programming and computer-based simulations, there are few works in the context of location analysis that are based on data analysis methods. The predictive power of various machine learning approaches (supervised regression, decision trees, support vector regression and pairwise ranking approach) to determine the ranking of potential localities for retail stores was tested in Karamshuk2013, concluding that geographic and mobility features strongly improve the results. The paper Chen2015 proposed the method for prediction of bike-sharing stations utilization using data on demographics, human activity, and area function as important factors influencing optimal placement of bike-sharing stations. Two-phase feature selection method to recognize useful features (derived from heterogeneous urban open data) for bike trip demand prediction is applied. In DSilva2018, it was demonstrated that similarity of urban neighborhoods and localities, and spatio-temporal features can be exploited to predict successfully the temporal activity patterns of new business venues using the k-nearest neighbor method and Gaussian processes.
In the electric vehicle domain, often, mobility data are used to estimate future demand. In De2015, GPS driving databases and data mining were used to plan EVCI. Real-world driving and parking events are examined to assess suitable locations for charging spots based on existing points of interest databases and a minimum-distance criterion. The efficient distribution of selected public charging points is investigated through discharge rate of EVs in tao2018data. The work Yang2017 introduced a data-driven optimization-based approach for the siting and sizing of electric taxi charging infrastructure in the city to minimize the infrastructure investments.
More recently, research works based on EV charging data have started to appear in the literature. A data-driven approach to extract useful information from EV charging events was suggested in Ref. Xydas2016. The proposed framework combines data pre-processing, data mining and fuzzy based models with real charging events data and weather data from three counties in the UK to characterize the charging demand of electric vehicles. Four well-known data mining algorithms, namely classification and regression trees (CART), Random forest algorithm (RFA), k-nearest neighbor (k-NN), and general chi-square automatic interaction detector (CHAID), were applied in Verma2015. The developed approach aims to identify and classify households with EVs by analyzing their energy consumption patterns. A data-driven statistical approach to extend the EVCI is noted in Pevec2018. This study suggested enriching EV charging datasets with contextual information (e.g. point of interest and driving distances between charging sites) to deploy new charging infrastructures. Hence, the benefits of geographical, demographic and economic data to optimal planning of EVCI have been acknowledged by some studies.
Realistic planning of EVCI can be achieved only if real-world data are available and accurate models are recognized. Several research studies in the field refer to the EVnetNL dataset, one of the biggest datasets available for the area of the Netherlands elaadnl. Contribution focusing on EV load forecasting by comparing time series (SARIMAX model) and machine learning approaches (Random Forest, Gradient Boosting Regression Trees) was presented in Ref. Buzna2019. Authors in Pevec2018 built EVCI utilization prediction models combining the EVnetNL dataset with business data, such as historical data about EV charging transactions and information about competitors in the market. The EVnetNL dataset was used to investigate and compare the performance of strategic and demand-driven rollout strategies for EVCI in the Netherlands. The obtained results in this study indicate that the proper rollout strategy depends on the maturity of the market (EV-adoption) and technology (battery capacity) Helmus2018. The EVnetNL dataset has been studied also in Develder2016 to analyze EV charging flexibility as demand response potential and in Lucas2019 to test the ability of regression algorithms to predict EV charging idle time. A set of eight indicators (e.g. EV energy demand, spatial density of EV chargers, etc.) was applied to the EVnetNL dataset to assess EVCI across countries Lucas2018. In addition, proposed indicators are used to assess the impact of relevant public policies on the rollout and utilization of EVCI. The work Flammini2019 analyzed the EVnetNL dataset of 400,000 EV charging transactions in the Netherlands for the year 2015 by using a weighted affine combination of beta distributions to estimate the multimodal probability distribution of charge time, connected time and idle time. By analyzing 390k transactions, the potential to shave the energy consumption peak was investigated, while using clustering techniques to categorize charging sessions by the arrival time, departure time and idle time Sadeghianpourhamami2018.
From the literature, it becomes apparent that planning of the EVCI requires an interdisciplinary approach that combines energy planning and management, economic and policy considerations, social science, geography, and data science. Therefore, our work is focused on an analytical framework that captures social, demographic, urban and transport characteristics to inform strategies for EVCI deployment and to optimize the utilization of the charging infrastructure. Many previous studies have optimized the placement of EVCI by minimizing the investments while ensuring a certain level of coverage of expected demand. The drawback of solely minimizing costs may lead to a design that is not corresponding to optimal usage patterns and hence will not generate sufficient return of investments. The popular sites might be associated with higher rollout costs but are more likely to return the investments and pay the maintenance costs or even be profitable. To our best knowledge, this is the first paper where a prediction model attempting to forecast the popularity of charging infrastructure is presented and validated. The main contributions of this paper are as follows:
The rest of the paper is organized as follows. More details on the EVnetNL dataset, used GIS data and selection of response variable are given in Section (ref). Section (ref) presents prediction methods, training, and validation of models. Obtained results including settings of parameters, measures of predictability and characteristics affecting the popularity of public charging infrastructure are documented in Section (ref). Discussion and summary of conclusions are provided in Section (ref).
Recently, a document suggesting unified terminology to be used in the electromobility field was published EVdefinitions. It defines a charging station as a physical object with one or more charging points sharing a common user identification interface. A charging point is an energy delivery device that might have one or several connectors, where only one can be used at the same time to charge an EV. A charging pool consists of one or multiple charging stations and the associated parking lots have one operator and a single address. Charging stations located close to each other have the same underlying geographical context and hence cannot be differentiated based on GIS data. In this paper, as an object of study, the charging pools are considered.
The EVnetNL dataset consists of more than one million charging transactions, performed on around 1 700 charging pools, distributed across the whole area of the Netherlands, by more than 50 000 EV users. Each transaction is initiated by plugging in and terminated by plugging out the EV. Each transaction is characterized by consumed energy, charging and connection time, unique identifier of a charging station and linked to EV user by RFID card. Data records span January 2012 to March 2016. The maximum available charging power at charging stations is 11 kW supporting merely slow charging. Transactions taking place in 2015 are considered in the analysis, as it is the last complete year in the dataset and the number of charging stations was already saturated. A more detailed description of the dataset can be found in Section (ref) of the Supplementary Information (SI) file.
Open GIS data were collected from various sources to model the urban context and human activities near charging pools. Brief descriptions of used datasets is provided in Table (ref). Datasets are available in various formats, hence, different predictor extraction techniques, described in Section (ref) of the SI file, were required. The extracted predictors have been pre-processed using workflow derived from kuhn2013applied,james2013introduction and detailed in Section (ref) of the SI file.
To properly inform the planning and deployment of charging infrastructure, there are many interdependent and often contradicting aspects that should be considered when quantifying the performance of charging pools:
A comprehensive set of performance indicators to compare EVCI rollout among continents, countries, and regions was proposed in Lucas2018. Having in mind different views of policymakers, municipalities, power system operators and charging infrastructure operators and considering the availability of data, we designed the following set of indicators to quantify the performance of charging pools:
The values of all performance indicators were calculated from the 2015 data. From the EVnetNL data, we found that if a charging pool has more than one connector, connectors were used only very rarely simultaneously. In the case of a charging pool with two connectors, this is because of construction reasons, in other cases this is probably due to a still relatively low number of electric vehicles in 2015. For this reason, the proposed indicators measure the performance of charging pools without discounting for the number of connectors. To analyze how well predictors derived from GIS data fit proposed performance indicators, we applied the ordinary least squares methods to each performance indicator separately. In Table (ref), we report values of the coefficient of determination, $R^2$. Results show that the highest $R^2$ value and hence the highest potential for data analysis is found for the popularity of charging pools (expressed by the unique number of RFID cards used to initiate charging). Therefore we limit further analyses presented in this paper to this performance indicator. To ensure transparency of presented analyses, as a supplementary material we publish together with the paper files containing values of investigated performance indicator and the matrix of predictors.
When planning the deployment of new infrastructure, often we need to select from a finite number of candidate sites (locations where it is feasible to install a charging infrastructure from the perspective of land ownership, supply with energy, potential to attract sufficient demand, etc.). In such a situation, it is not necessary to estimate the exact number of EV drivers attracted by a charging pool but to provide a ranking of candidate sites. For this reason, we reduce the problem that we address in this paper, to the problem of predicting whether a given candidate site belongs to the top rank candidate sites or not. This problem can be formalized as a classification problem. In the next section, we briefly introduce selected classification methods, logistic regression with an $\mathit{l}_1$ penalty, gradient boosted regression trees and random forests.
To describe classification methods, first we introduce a basic notation. We denote matrix of predictors as $\bold{X} \in \mathbb{R}^{n \times p}$. The matrix $\textbf{X}$ is formed by $i = 1, \dots, n$ observations $\textbf{x}_i = \{x_{i1}, \dots, x_{ip} \}$ (rows of $\textbf{X}$) and $j = 1, \dots, p$ predictor vectors $x_j~=~\{x_{1j}, \dots, x_{nj} \}^T$ (columns of $\textbf{X}$). A vector of response variables, we denote as $y = \{y_1, \dots, y_n\}$, with $y_i \in \{0, 1\}$ for $i = 1, \dots, n$, where $y_i = 1$ represents the situation in which the $i$-th charging pool belongs to the top $z\%$ of most popular stations and $y_i = 0$ otherwise.
Finally, from collected data we derived $p = 172$ predictors, each with $n=1271$ observations (charging pools). Although, $n > p$ holds, considering the number of observations, the number of predictors is relatively high. We assume that only a relatively small fraction of predictors plays an important role, i.e. we expect that the resulting model will be sparse. For this reason, we selected classification methods that can handle well sparsity Hastie_2009,Hastie_2015.
The logistic regression is a popular method for binary classification problems leading to solving a convex optimization problem Boyd_2004. The response is modeled as a random variable $Y \in\{0,1\}$ and the observation is modeled as a random variable $X \in \mathbb{R}^{p}$. The logistic model takes the form
where $\beta_0$ is the intercept and $\beta$ is the vector of regression coefficients. The maximum likelihood estimate of parameters $\beta_0$ and $\beta$ in Eq. ((ref)) is found by solving the optimization problem
Likewise, as in the LASSO method Tibshirani_1996, the logistic regression with $\mathit{l}_1$ penalty (LR-$\mathit{l}_1$) is obtained by adding $\mathit{l}_1$ regularization to the objective ((ref)), resulting in the optimization problem
that is solved for some $\lambda \geq 0$. The $\mathit{l}_1$ penalty enables to shrink less-informative coefficients $\beta$ to zero and thereby increases the simplicity and explanatory power of the model. Hyperparameter $\lambda$ in the objective function ((ref)) enables to set a trade-off between the quality of the fit and sparsity of the model.
Equation ((ref)) is also used to derive predictions. Estimated values of regression coefficients $\hat{\beta}_0$ and $\hat{\beta}$ together with the observation $\textbf{x}$ are plugged into the right hand side of Eq. ((ref)) and the resulting value is used as an estimate $\hat{y}$. Using hyperparameter $\theta$, the thresholding is used to transform $\hat{y}$ to the binary value. Hence, if $\hat{y} \geq \theta$ the prediction is $1$ and otherwise it is $0$.
Random Forests (RF) is a method based on the regression tree model Hastie_2009. The regression tree model predicts the target variables from ramified observations. More specifically, this method splits the training data into several subsets by applying conditions upon predictors. The model training represents a ramification process. When the tree is trained, branches grow from a single node, and every node determines a condition on a single predictor. A unique path is traced on the basis of the value of a single predictor, iteratively splitting the datasets into two children subsets. In order to determine the local optimal condition for the split, the Mean Squared Error (MSE) is minimized. RF allows diversifying the training of multiple regression trees, Breiman2001. Particularly, a number m of individual trees (i.e. the RF) is independently trained using a bagged (bootstrap aggregated) subset of the total training data. The m-th regression tree generates a prediction through the following equation:
where $\textbf{x}$ is an observation and ${w_i^{(m)}}(\textbf{x})$ is weight evaluated as follows:
with $L^{(m)}$ the leaf of the m-th tree individuated by $\textbf{x}$ , and $R_{L^{(m)}}$ the domain of this leaf. The function $1(\cdot)$ takes value 1 if the expression within the brackets is true and 0 otherwise. Thus, the m-th tree returns as a prediction the average value of all responses that belong to the same leaf node as the observation for which the prediction is made. Then, RF returns its prediction through a simple average of the predictions of the individual trees. Finally, to obtain binary prediction, thresholding is applied.
The Gradient boosting regression tree (GBRT) exploits regression trees, within a different framework. In this case, regression trees are fitted upon residuals of a weak learner in an iterative way. The fitting stops when the improvement brought by the last iteration is smaller than a fixed threshold friedman2001. At the generic iteration m, the prediction is given by a recursive equation:
where F is the prediction provided by the weak learner, $J^{(m)}$ is the number of terminal regions $R_1^{(m)}, \dots, R_{J^{(m)}}^{(m)}$ (each corresponding to one leaf node) and $\gamma_1^{(m)},\dots , \gamma_{J^{(m)}}^{(m)}$ are parameters estimated at each iteration m. The estimation of these parameters is accomplished by minimizing the MSE. The equation ((ref)) gives the prediction of the target variable. Finally, to obtain binary prediction, thresholding is applied.
In this section, the predictability of the popularity of charging pools is evaluated using various metrics. Since the LR-$\mathit{l}_1$ classification method is returning a sparse vector of regression coefficients, we evaluate the type and strength of the influence of predictors on the popularity as well.
A set of measures (see Table (ref)) was compiled to assess the performance of classification models from different perspectives. All measures can be calculated from the elements TP (true positives), TN (true negatives), FP (false positives) and FN (false negatives) of the confusion matrix kuhn2013applied.
The accuracy is the proportion of pools predicted correctly, the precision is proportion of correct predictions of popular pools and the sensitivity (also called true positive rate) is the proportion of popular pools predicted correctly. The fall--out (called also false positive rate) is a fraction of unpopular pools predicted incorrectly and it is used on the x-axis in the receiver operating characteristic (ROC) curve. The overall performance of a classifier can be evaluated by the area under the ROC curve (AUC) james2013introduction.
By analyzing the dependency between a measure of predictability and probabilistic threshold $\theta$ a suitable range for $\theta$ can be determined. When applying this approach to the accuracy for data with class imbalance (i.e. unequal number of $1$s and $0$s in the response vector) it may not be intuitive to determine models reaching good accuracy (e.g. for $25\%$ of $1$s in the response vectors, we reach value of accuracy $0.625$ by just randomly guessing $1$s with probability of $0.25$ and $0$s with probability of $0.75$). The F--score and Matthews’ correlation coefficient (MCC) are more balanced measures recommended for data with class imbalance kuhn2013applied,matthews1975comparison. The F--score combines precision and sensitivity, taking harmonic mean of both measures, i.e. it moves towards the lower of the values sasaki2007truth. By definition, if any of the values in the parentheses in the denominator of the equation defining MCC is $0$ (see Table (ref)), the MCC is set to $0$. For models predicting better (worse) than a random model, the MCC is positive (negative). Models as skilled as a random guess have MCC equal to $0$. The MCC equals to $1$ ($-1$) if all the observations are predicted correctly (wrongly).
A radius of the buffer, representing vicinity of a pool, was chosen from the set $\{100, 150, 200, 250, 300, 350, 400, 450, 500\}$ meters. This range is the distance drivers typically walk from the parking place to their destination Waerden_2015. The dependency between the popularity of charging pools and predictors was fitted by the least squares model for all values in the set. The highest value of $R^2$ was found for the radius of $350$ meters and selected for further analyses. In experiments, training and testing sets are assigned $80\%$ and $20\%$, respectively, using stratified sampling with respect to the response variable kuhn2013applied. To gain more reliable results and conclusions, we evaluate variability of calculated measures by analyzing outputs of 100 models, trained on 100 different splits into training and testing sets. While evaluating predictability measures, we varied threshold $\theta$ in the range from 0 to 0.99, in steps of 0.01. The hyperparameter $\lambda$ in Eq. ((ref)), was found by k-fold cross validation, where the stratified sampling was used to split data into $k=10$ folds preserving the class distribution of the response variable. Considering values $10^{-4 + i*0.015}$, for $i \in\{0, ..., 200\}$, $\lambda$ is assigned the value corresponding to the largest AUC value friedman2010regularization. Similarly, when growing decision trees, the k-fold cross validation was applied to set the number of learning cycles, the learn rate for shrinkage, the minimum size of leaves, and the maximum number of splits. The values of predictability measures were obtained by applying the trained model to the testing dataset. The response vector was encoded into a binary format by setting $25\%$ of the elements corresponding to top charging pools ranked by the popularity to value $1$ and the rest to $0$. We did also experiments with values of $15\%$, $20\%$, $30\%$ and $35\%$, however, very similar conclusions could be drawn from the results and we do not report them. In computations, we used $\mathit{l}_1$-regularized logistic regression implemented by the R package glmnet friedman2010regularization. GBRT and RF were implemented in MATLAB environment within Statistic and Machine Learning Toolbox.
In Figure (ref), the mean accuracy, precision and sensitivity are shown for all three methods. To facilitate evaluation of the quality of predictions, we consider a null model predicting popular stations randomly with probability of $0.25$. It can be easily shown that considered measures for the null model take the functional forms presented in Table (ref) and displayed in Figures (ref) and (ref) by thin lines.
All three methods outperform the null model, in accuracy and precision measures in the whole range of $\theta$. It is unavoidable that the sensitivity is decreasing as the threshold $\theta$ grows. For $\theta$ values greater than $0.5$ the sensitivity falls below 0.5, which results in low applicability of predictions. As $\theta$ grows, the sensitivity can reach value zero, even for an ideal model, if the threshold $\theta$ is already too high. If $\theta$ is larger than 0.5, for the decision tree methods we find the sensitivity to be smaller than the null model confirming that such thresholds are too high.
In the literature, various approaches on how to select the value of the threshold $\theta$ and hence to find a reasonable trade-off between accuracy, precision and sensitivity can be found kuhn2013applied. We selected two metrics, MCC and F--score, which are evaluated in Figure (ref). Both measures reach one single maximum, which is hence also the global maximum. Values of the threshold where the maxima are achieved we denote as $\theta_{MCC_{max}}$ and $\theta_{F\textendash score_{max}}$. Values of the accuracy, precision and sensitivity for $\theta_{MCC_{max}}$ and $\theta_{F\textendash score_{max}}$ are reported in Table (ref).
When decisions about the extension of the existing charging infrastructure are taken, many often contradicting factors are taken into account. Alternatively, stakeholders can select a threshold $\theta$ based on their expectations and attitudes towards risk. The lower the $\theta$, the more likely is identification of popular locations with the drawback of increased risk of placing charging pools into unpopular areas. In opposite, the higher $\theta$, we identify the popular charging pools with higher assurance, with the drawback of overlooking potentially popular locations. According to observed values of measures, we recommend considering the threshold $\theta$ within the range from 0.3 to 0.45, where both precision and sensitivity are relatively high.
Comparison with the null model and values of measures in Table (ref) indicate that urban area and characteristics of charging pools contain some predictive power for the popularity. What values of measure justify good quality of models typically depend on the application domain james2013introduction. Considering data analyses concerning the human choice in similar domains such as for example bike sharing applications zhang2016bicycle, chen2017understanding, the obtained values of the accuracy exceeding value 0.8, while both precision and sensitivity are larger than 0.65, can be considered as favorable.
Simple inspection of Figures (ref)- (ref) indicates, that all three methods provide similar results. To evaluate, whether the results are statistically distinguishable, we test the differences in AUC ensembles, i.e. areas under the ROC curves. The ROC curves corresponding to 100 different training and testing dataset splits are shown in Figure (ref). In all statistical tests we set the significance level $\alpha = 0.01$.
In the first step, the equality of AUC variances between all pairs of methods was tested using the F-test. Statistically significant difference was identified only for LR-$\mathit{l}_1$ and RF methods, the first having less variance ($p = 0.0017$). To compare mean values of AUCs, for the case when the variances were not found to be significantly different, the t-test was used otherwise, the Welch t-test welch1947generalization. The mean AUC of LR-$\mathit{l}_1$ was larger than RF ($p = 0.0094$), and not different from GBRT ($p = 0.0379$). The mean AUCs of RF and GBRT were not significantly different ($p = 0.5091$). Hence, according to AUC values, the method LR-$\mathit{l}_1$ is more stable than RF, and outperforms the method in mean values of AUC. On the significance level $\alpha = 0.05$, LR-$\mathit{l}_1$ method has a significantly larger mean value of AUCs than both, the GBRT and RF methods.
The shrinkage nature of $\mathit{l}_1$ penalty in the LR-$\mathit{l}_1$ method provides better predictions on testing data and variable selection functionality. This gives an advantage compared to decision tree methods, that yielded clumsy models, involving 169 predictors on average, making them hard to interpret. Due to this reason, we interpret the results for the LR-$\mathit{l}_1$ method only.
The values $\hat{\beta}$ of the LR-$\mathit{l}_1$ model might be sensitive to the used sample of training data, hence an analysis of the results robustness is necessary. The conventional solution is to evaluate the $p$-values of coefficients estimated by statistical methods. The problem to calculate $p$-values for LR-$\mathit{l}_1$ model is difficult due to the adaptive nature of the estimation procedure tibshirani2015statistical. Therefore, to capture the stochasticity in coefficients $\hat{\beta}$, we estimate their distributions by sampling the dataset 500 times using bootstrap james2013introduction. A model is fitted to each bootstrapped dataset using stratified cross-validation. To facilitate a comparison of the impact of predictors, we standardize each element of $\hat{\beta}$ by dividing it with the sample standard deviation of the corresponding predictor.
The group of selected predictors depends on the training data sample. In Figure (ref), sampled distributions of coefficients that have been selected most frequently (at least by 90% out of 500 models) are displayed. The bar plot presents the percentage of models where the coefficients were equal to zero.
The impact of a predictor on the response increases with the absolute value of the corresponding coefficient kuhn2013applied. Positive (negative) sign of a coefficient indicates increasing (decreasing) impact of the predictor. A way how to quantify the significance of coefficient $\hat{\beta}$ is to assess the likelihood that the coefficient is different from zero. Analysis of the distribution of coefficients $\hat{\beta}$ allows us to make such assessments and to conclude how certain is the positive or negative influence of predictors on the response variable.
In summary, selected predictors can be categorized into three groups: the function of the geographic area constituting the vicinity of a charging pool, characteristics of the population living in this area and properties of charging pools. From the perspective of the geographic area, the most important predictors are the number of wholesale businesses, shops, hotels, restaurants and catering businesses, areas with recreational inland water, sports fields and roads, all having a positive impact. The minimal distance to financial, cultural, and transportation OpenStreetMap (OSM) amenities have negative coefficients, meaning that large minimal distance is decreasing the popularity of charging pools. Hence, if these facilities are found in the proximity of charging pools, they have a positive impact on the popularity.
In opposite, the residential areas, the areas with non-commercial ornamental and vegetable cultivation and the presence of OSM amenities related to households tend to lower the popularity of charging pools. These findings are well aligned with the intuition that in residential areas the charging pools are visited by more homogeneous groups of users than in the more crowded urban areas. Similarly, the most likely explanation of the negative impact of business and industrial areas is work charging Sadeghianpourhamami_2018, i.e. either charging of a fleet of company cars or regular use of charging pools by (a small group of) employees commuting to work.
A notable population group living near popular charging pools is working elderly people between 65 and 74 years old. In opposite, popularity is negatively correlated with areas inhabited by the population working in the mining, manufacturing or construction sector as well as persons who depend on social assistance. Thus, these results suggest that the economic prosperity of the population in the vicinity is affecting the visiting patterns of charging pools. The popular charging pools are more likely to be deployed following strategical rollout and have larger maximal power and more connectors. The negative influence of geographic longitude can be explained by the geography of the Netherlands, while the western part of the country is more urbanized and we can find here the majority of large Dutch cities.
This study demonstrates the ability of classification methods to predict popular locations of charging pools from various GIS and large-scale charging infrastructure data. Moreover, we evaluate the impact and significance of factors affecting the popularity of charging infrastructure. Predicting the popularity of charging pools is of utmost importance for matching EV requirements driven by social habits with energy requirements related to electrical network configuration. Main conclusions derived from the data analysis are the following:
The presented results are limited by the low utilization of some charging pools, which can be attributed to the low penetration of electric vehicles in some areas in the Netherlands. Thanks to the rapid growth of EVs, the utilization of charging pools is expected to increase and potentially mitigate this limitation. A difficulty often present in GIS data is the interdependence of factors, expressed as collinearity, causing nontrivial problems when interpreting the impact of individual factors. There is no generally accepted approach to address this problem. We minimize the chances that collinearity affects the results by removing highly correlated factors. Nevertheless, predictors that we present as influential and significant, should be taken into account with some care. Typically, the uncaptured stochasticity of models can be attributed to missing data. We assume that more detailed mobility data, e.g. GPS and floating car data could improve the results. Geographically, our study is focused on the area of the Netherlands, which might impose some limitations when transferring models and conclusions to other countries. Although, we expect similar results for comparable geographic and demographic contexts.
This study opens several directions. For instance, future research could explore possibilities how to design prediction models for other performance indicators of charging pools, how to efficiently downscale prediction models to a level of a region or a city or how to improve predictions by customizing models to specific classes of charging pools. Another challenge is the application of regression approaches that could successfully predict the values of performance indicators.
This work was supported by the research grants: VEGA 1/0089/19 ”Data analysis methods and decisions support tools for service systems supporting electric vehicles”, VEGA 1/342/18 ”Optimal dimensioning of service systems”, APVV-15-0179 ”Reliability of emergency systems on infrastructure with uncertain functionality of critical elements”, Operational Program Research and Innovation in frame of the project: ICTproducts for intelligent systems communication, code ITMS2014+ 313011T413, co-financed by the European Regional Development Fund and by the Slovak Research and Development Agency under the contract no. SK-IL-RD-18-005. We thank \v{L}udmila J\'{a}no\v{s}\'{i}kov\'{a} and Luca Lena Jansen for valuable comments and suggestions.
CRediT (Contributor Roles Taxonomy) has been applied to describe contributions of authors. Milan Straka: Data curation, Methodology, Formal analysis, Investigation, Validation, Visualization, Software, Writing-original draft; Pasquale De Falco: Formal analysis, Methodology, Investigation, Software, Validation, Writing-original draft; Gabriela Ferruzzi: Validation; Writing original draft; Writing-review and editing; Daniela Proto: Validation; Writing original draft; Writing-review and editing; Gijs van der Poel: Conceptualization, Resources, Writing-review and editing; Shahab Khormali: Resources, Writing-original draft, Writing-review and editing; \textbf{\v{L}ubo\v{s} Buzna:} Conceptualization, Funding acquisition, Methodology, Resources, Supervision, Writing-review and editing.
Competing financial interests: The authors declare no competing financial interests.
\setcounter{section}{0}