The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
76,072 characters
Predicting popularity of EV charging infrastructure from GIS data
\begin{frontmatter}
\title{Predicting popularity of EV charging infrastructure from GIS data\tnoteref{label1}}
\author[1]{Milan~Straka}
\ead{[email removed]}
\author[2]{Pasquale~De~Falco}
\author[3]{Gabriella~Ferruzzi}
\author[3]{Daniela~Proto}
\author[4]{Gijs~van~der~Poel}
\author[1]{Shahab~Khormali}
\author[1]{\v{L}ubo\v{s} Buzna}
\cortext[cor1]{Corresponding author.}
\address[1]{University of \v{Z}ilina, Univerzitn\'{a} 8215/1, \v{Z}ilina, Slovakia}
\address[2]{Department of Engineering, University of Napoli Parthenope, Naples, Italy}
\address[3]{Department of Information Technology and Electrical Engineering, University of Napoli Federico II, Naples, Italy}
\address[4]{ElaadNL, Utrechtseweg 310 (bld. 42B), 6812 AR Arnhem (GL), The Netherlands}
\begin{abstract}
The availability of charging infrastructure is essential for large-scale adoption of electric vehicles (EV). Charging patterns and the utilization of infrastructure have consequences not only for the energy demand, loading local power grids but influence the economic returns, parking policies and further adoption of EVs. We develop a data-driven approach that is exploiting predictors compiled from GIS data describing the urban context and urban activities near charging infrastructure to explore correlations with a comprehensive set of indicators measuring the performance of charging infrastructure. The best fit was identified for the size of the unique group of visitors (popularity) attracted by the charging infrastructure. Consecutively, charging infrastructure is ranked by popularity. The question of whether or not a given charging spot belongs to the top tier is posed as a binary classification problem and predictive performance of logistic regression regularized with an $\mathit{l}_1$ penalty, random forests and gradient boosted regression trees is evaluated. Obtained results indicate that the collected predictors contain information that can be used to predict the popularity of charging infrastructure. The significance of predictors and how they are linked with the popularity are explored as well. The proposed methodology can be used to inform charging infrastructure deployment strategies.
\end{abstract}
\begin{graphicalabstract}
\includegraphics[width=1.0\textwidth]{graphical_abstract}
\end{graphicalabstract}
\begin{highlights}
\item Economic and urban context predict popularity of charging infrastructure.
\item $l_1$-regularized logistic regression modestly outperforms decision trees.
\item Popularity is positively linked with charging infrastructure specs, hotels, food and financial services.
\item Popularity is negatively linked with residential areas and distance from frequently visited venues.
\end{highlights}
\begin{keyword}
electric vehicles \sep deployment policies for charging infrastructure \sep demand and popularity \sep prediction models \sep economic and spatial analysis
\end{keyword}
\end{frontmatter}
\section{Introduction}
\label{sec:intro}
Energy consumption has increased outstandingly in the last years and it will continue to be a significant global challenge. The largest portion of the total energy consumption is in the transportation sector where high economic and population growth causes a rapid increase in energy demand with excessive CO\textsubscript{2} emissions and energy crisis~\cite{EIA2013}. To mitigate the impact of the emissions of greenhouse gases and to increase energy security, automotive powertrain electrification in transportation sector could play an important role~\cite{Tali_2013, wilbanks2010}. This is confirmed by the fact that several countries have announced an ambition to stop selling vehicles fueled by diesel or petrol. For example, the UK set as the target year 2040. The EU has taken a decisive step forward in implementing the EU’s commitments under the Paris Agreement for a binding domestic CO\textsubscript{2} reduction of at least 40\% until 2030~\cite{JRC_web}. Electric vehicles (EVs), as a key element of clean and green travel mode, are spreading all over the world rapidly. For instance, an ambitious target of having over 20 million EVs on the roads by 2020 has been set by the U.S. Department of Energy~\cite{EIA2013}. However, the adoption of EVs on a large scale is supposed to bring both challenges and opportunities from technical and economic points of view~\cite{ajanovic2016dissemination,Flammini2017}.
\subsection{Motivation}
\label{sec:motivation}
In order to boost EV popularity, many challenges need to be addressed.
The chicken-egg problem in the form of charging infrastructure versus EV adoption has been recognized as an important challenge restraining the growing EV ecosystem \cite{van2013data}. Currently, insufficient charging infrastructure is a significant factor that prevents larger penetration of EVs~\cite{Coffman2017}. The drivers are hesitating to buy an EV if there is not sufficient charging infrastructure, and similarly charging infrastructure operators do not invest while the number of EVs is low and not profitable. In recent years, demand-driven and strategic rollouts (i.e. strategically covering the geographic space by chargers) have been applied~\cite{Helmus2018}. The opening of a new public charging infrastructure involves the estimation of the visitation patterns to ensure as high as possible utilization to justify the allocated resources.
Prior to opening a new public charging infrastructure, the expected utilization should be estimated in order to justify the allocated resources. Hence, one way how to support decision making in this area is to develop data-driven approaches with a predictive power.
\subsection{Previous relevant work}
\label{sec:previous_relevant_work}
Different methods, such as mathematical programming, computer-based simulation and statistical learning, have been proposed to deploy EV charging infrastructure (EVCI). Mathematical programming models for the optimal location of EVCI have been proposed in several papers. The methodologies are based on the minimization of objective functions, addressing infrastructure development costs~\cite{Liu2012, Sadeghi2014}, social costs~\cite{ren2019location}, driving range and driver habits~\cite{Gonzalez2014,Dong2014,He2018}, traffic flow data~\cite{Efthymiou2017}, unmet demand~\cite{Chen2013}, and quality of service~\cite{Davidov2017}. In \cite{GE2012} was addressed the problem considering road network and distribution system network capacity constraints. The work of Asamer et al.~\cite{Asamer2016} provides a formulation of the optimal location problem for a specific category of vehicles (taxi service). A comprehensive review of different optimization techniques for EVCI can be found~\cite{Rahman2016}.
In contrast to the large number of studies that applied mathematical programming and computer-based simulations, there are few works in the context of location analysis that are based on data analysis methods. The predictive power of various machine learning approaches (supervised regression, decision trees, support vector regression and pairwise ranking approach) to determine the ranking of potential localities for retail stores was tested in~\cite{Karamshuk2013}, concluding that geographic and mobility features strongly improve the results. The paper~\cite{Chen2015} proposed the method for prediction of bike-sharing stations utilization using data on demographics, human activity, and area function as important factors influencing optimal placement of bike-sharing stations. Two-phase feature selection method to recognize useful features (derived from heterogeneous urban open data) for bike trip demand prediction is applied. In~\cite{DSilva2018}, it was demonstrated that similarity of urban neighborhoods and localities, and spatio-temporal features can be exploited to predict successfully the temporal activity patterns of new business venues using the k-nearest neighbor method and Gaussian processes.
In the electric vehicle domain, often, mobility data are used to estimate future demand. In~\cite{De2015}, GPS driving databases and data mining were used to plan EVCI. Real-world driving and parking events are examined to assess suitable locations for charging spots based on existing points of interest databases and a minimum-distance criterion. The efficient distribution of selected public charging points is investigated through discharge rate of EVs in~\cite{tao2018data}. The work~\cite{Yang2017} introduced a data-driven optimization-based approach for the siting and sizing of electric taxi charging infrastructure in the city to minimize the infrastructure investments.
More recently, research works based on EV charging data have started to appear in the literature. A data-driven approach to extract useful information from EV charging events was suggested in Ref.~\cite{Xydas2016}. The proposed framework combines data pre-processing, data mining and fuzzy based models with real charging events data and weather data from three counties in the UK to characterize the charging demand of electric vehicles. Four well-known data mining algorithms, namely classification and regression trees (CART), Random forest algorithm (RFA), k-nearest neighbor (k-NN), and general chi-square automatic interaction detector (CHAID), were applied in~\cite{Verma2015}. The developed approach aims to identify and classify households with EVs by analyzing their energy consumption patterns. A data-driven statistical approach to extend the EVCI is noted in~\cite{Pevec2018}. This study suggested enriching EV charging datasets with contextual information (e.g. point of interest and driving distances between charging sites) to deploy new charging infrastructures. Hence, the benefits of geographical, demographic and economic data to optimal planning of EVCI have been acknowledged by some studies.
Realistic planning of EVCI can be achieved only if real-world data are available and accurate models are recognized. Several research studies in the field refer to the EVnetNL dataset, one of the biggest datasets available for the area of the Netherlands~\cite{elaadnl}. Contribution focusing on EV load forecasting by comparing time series (SARIMAX model) and machine learning approaches (Random Forest, Gradient Boosting Regression Trees) was presented in Ref.~\cite{Buzna2019}. Authors in~\cite{Pevec2018} built EVCI utilization prediction models combining the EVnetNL dataset with business data, such as historical data about EV charging transactions and information about competitors in the market. The EVnetNL dataset was used to investigate and compare the performance of strategic and demand-driven rollout strategies for EVCI in the Netherlands. The obtained results in this study indicate that the proper rollout strategy depends on the maturity of the market (EV-adoption) and technology (battery capacity)~\cite{Helmus2018}. The EVnetNL dataset has been studied also in~\cite{Develder2016} to analyze EV charging flexibility as demand response potential and in~\cite{Lucas2019} to test the ability of regression algorithms to predict EV charging idle time. A set of eight indicators (e.g. EV energy demand, spatial density of EV chargers, etc.) was applied to the EVnetNL dataset to assess EVCI across countries~\cite{Lucas2018}. In addition, proposed indicators are used to assess the impact of relevant public policies on the rollout and utilization of EVCI. The work~\cite{Flammini2019} analyzed the EVnetNL dataset of 400,000 EV charging transactions in the Netherlands for the year 2015 by using a weighted affine combination of beta distributions to estimate the multimodal probability distribution of charge time, connected time and idle time. By analyzing 390k transactions, the potential to shave the energy consumption peak was investigated, while using clustering techniques to categorize charging sessions by the arrival time, departure time and idle time~\cite{Sadeghianpourhamami2018}.
\subsection{Our contribution}
\label{sec:our_contribution}
From the literature, it becomes apparent that planning of the EVCI requires an interdisciplinary approach that combines energy planning and management, economic and policy considerations, social science, geography, and data science. Therefore, our work is focused on an analytical framework that captures social, demographic, urban and transport characteristics to inform strategies for EVCI deployment and to optimize the utilization of the charging infrastructure. Many previous studies have optimized the placement of EVCI by minimizing the investments while ensuring a certain level of coverage of expected demand. The drawback of solely minimizing costs may lead to a design that
is not corresponding to optimal usage patterns and hence will not generate sufficient return of investments. The popular sites might be associated with higher rollout costs but are more likely to return the investments and pay the maintenance costs or even be profitable. To our best knowledge, this is the first paper where a prediction model attempting to forecast the popularity of charging infrastructure is presented and validated. The main contributions of this paper are as follows:
\begin{itemize}
\item We summarize goals pursued by stakeholders involved in the charging infrastructures development while selecting and describing the set of performance indicators approximating their intentions.
\item We provide an analytical framework that captures social, demographic, urban and transport characteristics and the availability of charging opportunities of urban neighborhoods surrounding the charging infrastructure.
\item We evaluate the predictive power of three prediction models: logistic regression regularized with an $l_1$ penalty, random forests and gradient boosted regression trees.
\item From results, we evaluate the significance of factors affecting the popularity of charging infrastructure.
\end{itemize}
The rest of the paper is organized as follows. More details on the EVnetNL dataset, used GIS data and selection of response variable are given in Section~\ref{sec:materials}. Section~\ref{sec:methods} presents prediction methods, training, and validation of models. Obtained results including settings of parameters, measures of predictability and characteristics affecting the popularity of public charging infrastructure are documented in Section~\ref{sec:results}. Discussion and summary of conclusions are provided in Section~\ref{sec:discussion_and_conslusions}.
\section{Materials}
\label{sec:materials}
\subsection{Charging pools}
\label{sec:charging_pools}
Recently, a document suggesting unified terminology to be used in the electromobility field was published~\cite{EVdefinitions}. It defines a charging station as a physical object with one or more charging points sharing a common user identification interface. A charging point is an energy delivery device that might have one or several connectors, where only one can be used at the same time to charge an EV. A charging pool consists of one or multiple charging stations and the associated parking lots have one operator and a single address. Charging stations located close to each other have the same underlying geographical context and hence cannot be differentiated based on GIS data. In this paper, as an object of study, the charging pools are considered.
\subsection{EVnetNL dataset}
The EVnetNL dataset consists of more than one million charging transactions, performed on around 1~700 charging pools, distributed across the whole area of the Netherlands, by more than 50~000 EV users. Each transaction is initiated by plugging in and terminated by plugging out the EV. Each transaction is characterized by consumed energy, charging and connection time, unique identifier of a charging station and linked to EV user by RFID card. Data records span January 2012 to March 2016. The maximum available charging power at charging stations is 11~kW supporting merely slow charging. Transactions taking place in 2015 are considered in the analysis, as it is the last complete year in the dataset and the number of charging stations was already saturated. A more detailed description of the dataset can be found in Section~\ref{secSI:EVnetNL_dataset} of the Supplementary Information (SI) file.
\subsection{GIS datasets}
Open GIS data were collected from various sources to model the urban context and human activities near charging pools. Brief descriptions of used datasets is provided in Table~\ref{tab:GISdata}. Datasets are available in various formats, hence, different predictor extraction techniques, described in Section~\ref{secSI:EVnetNL_dataset} of the SI file, were required. The extracted predictors have been pre-processed using workflow derived from~\cite{kuhn2013applied,james2013introduction} and detailed in Section~\ref{secSI:data_pre_processing} of the SI file.
\begin{longtable}[h]{p{.2\textwidth} p{.7\textwidth} p{.1\textwidth} }
\hline
Dataset & Brief description & Source\\
\hline
Population cores & Detailed population data organized by morphologically continuous areas. & \cite{popcores} \\
Neighborhoods & Population data aggregated to neighborhoods. & \cite{neighbpop}\\
Land use & Land use data modeled by high resolution heterogeneous polygons and divided into 25 categories. & \cite{lccbs} \\
Energy & Aggregated gas and electricity consumption of companies and households estimated per neighborhoods. & \cite{energyatlas} \\
Liveability & General index describing quality of living in 5 categories (housing, residents, services, safety, and living environment) at the level of neighborhoods. & \cite{liveability} \\
Traffic flows & Database of traffic volumes of cars, buses and trucks on individual roads. & \cite{traffic_flows}\\
LandScan & Ambient population density in the raster format. & \cite{landscan}\\
OpenStreetMap & OpenStreetMap amenities taking form of the point data. & \cite{osm} \\
Charging pools 2015 & Available locations of charging pools in 2015. & \cite{ocm,opp} \\
\hline
\caption{Overview of collected GIS datasets.}
\label{tab:GISdata}
\end{longtable}
\subsection{Selection of response variable: The popularity of charging pools}
\label{sec:response_variable}
To properly inform the planning and deployment of charging infrastructure, there are many interdependent and often contradicting aspects that should be considered when quantifying the performance of charging pools:
\begin{itemize}
\item Being motivated by the need to innovate the road transport that is based on fossil fuels, from the perspective of governments, policymakers, and municipalities, one of the main goals when developing charging infrastructure is to stimulate wide use of EVs and to invest public resources efficiently and fairly.
\item At present, the main concern of grid operators is the stability of supply systems and seamless integrations of renewable energy sources. Technologies, such as smart charging, should help to harvest the potential of EVs in reaching this goal. This requires alignment between the presence of energy produced from renewables and the charging behavior.
\item When building or extending the network, charging infrastructure operators are seeking a profit. This requires placing charging pools in locations that can generate sufficient revenues.
\end{itemize}
A comprehensive set of performance indicators to compare EVCI rollout among continents, countries, and regions was proposed in~\cite{Lucas2018}. Having in mind different views of policymakers, municipalities, power system operators and charging infrastructure operators and considering the availability of data, we designed the following set of indicators to quantify the performance of charging pools:
\begin{itemize}
\item \textit{Consumed energy:} Dependent on the subscription program, EV drivers either pay regularly a fixed amount and have unlimited access to charging services or they are charged a fee when connecting the vehicle and pay a certain rate per unit of consumed energy. Hence, consumed energy on a charging pool is a proxy of profitability, moreover, it also indicates how difficult it is to integrate a charging pool to the power grid.
\item \textit{Number of charging transactions:} The higher the number of EVs visiting a charging pool, the higher the potential profit. On the one hand, a high number of transactions is loading the supply network more, one the other hand it can create more opportunities for smart charging.
\item \textit{Popularity of a charging pool:} The ability of a charging pool to attract a large group of EV drivers can be approximated by the number of unique RFID cards used on a pool. Popular pools can be more robust concerning random fluctuations in the usage compared to charging pools that are highly used but only by a small group of EV drivers. Moreover, public investments into popular charging pools can be considered as socially fairer.
\item \textit{Charging time:} Complementary information about the usage of a charging pool is provided by the overall charging time (i.e. the time that is used to transfer energy between charging infrastructure and vehicle). Long charging time indicates a high utilization of a charging pool, but it can also indicate that the rated power of a charging pool should be increased.
\item \textit{Charging ratio:} An issue occurs when EV drivers tend to leave the vehicle at the charging pool for a much longer time than is required for charging. This behavior can be motivated by free parking at charging pools in cities, and it limits the access to charging spots for other EV users. Charging ratio is a number between $0$ and $1$ that is calculated by dividing the length of the time intervals that vehicles are charging with the length of the time intervals vehicles are connected to a given charging pool. On one hand, a charging pool that has a high charging ratio can reach higher profit as it is better utilized by EV drivers. On the other hand, a low charging ratio means a higher potential for smart charging.
\item \textit{Use-time ratio:} If the occupancy of a pool is high, other EV drivers are often forced to search for charging alternatives, which is decreasing the perceived quality of service provided by the charging pools. A simple measure, estimating the occupancy of a charging pool is the use-time ratio~\cite{Lucas2018}, which is calculated by dividing the length of the time period when the pool was occupied by the length of the observed period.
\item \textit{Energy ratio:} How much a charging pool is loading the supply network, can be estimated by the energy ratio~\cite{Lucas2018}, which is calculated by dividing the consumed energy by the rated energy (i.e. energy that would be consumed by the charging pool if working at the maximum capacity for the whole observed period). This indicator may inform power grid operators about the possible supply margin that could be eventually used by other consumers and it is one of the indicators that can help charging infrastructure operators to understand better the utilization of charging pools.
\end{itemize}
\begin{table}[ht]
\centering
\begin{tabular}{ll}
\hline
Charging pool performance indicator & $R^2$ \\
\hline
Consumed energy [kWh] & 0.44\\
Number of charging transactions & 0.50\\
Popularity (unique number of RFID cards) & 0.60\\
Charging time [hours] & 0.48 \\
Charging ratio (charging time divided by connection time) & 0.40 \\
Use-time ratio (charging time divided by overall operation time) & 0.42 \\
Energy ratio (consumed energy divided by rated energy) & 0.38 \\
\hline
\end{tabular}
\caption{Assessment of performance indicators using predictors characterizing urban context and human activities near charging pools. The values of all performance indicators were calculated from the 2015 data. The coefficient of determination, $R^2$, was obtained by the ordinary least squares method.}
\label{tab:R2regResponses}
\end{table}
The values of all performance indicators were calculated from the 2015 data. From the EVnetNL data, we found that if a charging pool has more than one connector, connectors were used only very rarely simultaneously. In the case of a charging pool with two connectors, this is because of construction reasons, in other cases this is probably due to a still relatively low number of electric vehicles in 2015. For this reason, the proposed indicators measure the performance of charging pools without discounting for the number of connectors. To analyze how well predictors derived from GIS data fit proposed performance indicators, we applied the ordinary least squares methods to each performance indicator separately. In Table~\ref{tab:R2regResponses}, we report values of the coefficient of determination, $R^2$. Results show that the highest $R^2$ value and hence the highest potential for data analysis is found for the popularity of charging pools (expressed by the unique number of RFID cards used to initiate charging). Therefore we limit further analyses presented in this paper to this performance indicator. To ensure transparency of presented analyses, as a supplementary material we publish together with the paper files containing values of investigated performance indicator and the matrix of predictors.
When planning the deployment of new infrastructure, often we need to select from a finite number of candidate sites (locations where it is feasible to install a charging infrastructure from the perspective of land ownership, supply with energy, potential to attract sufficient demand, etc.). In such a situation, it is not necessary to estimate the exact number of EV drivers attracted by a charging pool but to provide a ranking of candidate sites. For this reason, we reduce the problem that we address in this paper, to the problem of predicting whether a given candidate site belongs to the top rank candidate sites or not. This problem can be formalized as a classification problem. In the next section, we briefly introduce selected classification methods, logistic regression with an $\mathit{l}_1$ penalty, gradient boosted regression trees and random forests.
\section{Methods}
\label{sec:methods}
To describe classification methods, first we introduce a basic notation. We denote matrix of predictors as $\bold{X} \in \mathbb{R}^{n \times p}$. The matrix $\textbf{X}$ is formed by $i = 1, \dots, n$ observations $\textbf{x}_i = \{x_{i1}, \dots, x_{ip} \}$ (rows of $\textbf{X}$) and $j = 1, \dots, p$ predictor vectors $x_j~=~\{x_{1j}, \dots, x_{nj} \}^T$ (columns of $\textbf{X}$). A vector of response variables, we denote as $y = \{y_1, \dots, y_n\}$, with $y_i \in \{0, 1\}$ for $i = 1, \dots, n$, where $y_i = 1$ represents the situation in which the $i$-th charging pool belongs to the top $z\%$ of most popular stations and $y_i = 0$ otherwise.
Finally, from collected data we derived $p = 172$ predictors, each with $n=1271$ observations (charging pools). Although, $n > p$ holds, considering the number of observations, the number of predictors is relatively high. We assume that only a relatively small fraction of predictors plays an important role, i.e. we expect that the resulting model will be sparse. For this reason, we selected classification methods that can handle well sparsity~\cite{Hastie_2009,Hastie_2015}.
\subsection{Logistic regression with $\mathit{l}_1$ penalty}
\label{sec:l1_logistic_regression}
The logistic regression is a popular method for binary classification problems leading to solving a convex optimization problem~\cite{Boyd_2004}. The response is modeled as a random variable $Y \in\{0,1\}$ and the observation is modeled as a random variable $X \in \mathbb{R}^{p}$. The logistic model takes the form
\begin{equation}
P(Y = 1|X = x) = \frac{1}{1+e^{-(\beta_0 + \beta^T x)}},
\label{eq:log_reg}
\end{equation}
where $\beta_0$ is the intercept and $\beta$ is the vector of regression coefficients. The maximum likelihood estimate of parameters $\beta_0$ and $\beta$ in Eq.~(\ref{eq:log_reg}) is found by solving the optimization problem
\begin{align}
\underset{\beta_0, \beta}{\text{maximize }} & \frac{1}{n}\sum_{i=1}^{n}\left \{{y_i(\beta_0 + \beta^{T} \textbf{x}_i)-\log{(1+e^{\beta_0 + \beta^T \textbf{x}_i})}}\right\}.
\label{eq:log_reg_obj}
\end{align}
Likewise, as in the LASSO method~\cite{Tibshirani_1996}, the logistic regression with $\mathit{l}_1$ penalty (LR-$\mathit{l}_1$) is obtained by adding $\mathit{l}_1$ regularization to the objective~(\ref{eq:log_reg_obj}), resulting in the optimization problem
\begin{align}
\underset{\beta_0, \beta}{\text{maximize }} & \frac{1}{n}\sum_{i=1}^{n}\left \{{y_i(\beta_0 + \beta^{T} \textbf{x}_i)-\log{(1+e^{\beta_0 + \beta^T \textbf{x}_i})}}\right\} + \lambda \left\|{\beta}\right\|_1,
\label{eq:log_reg_lasso_obj}
\end{align}
that is solved for some $\lambda \geq 0$. The $\mathit{l}_1$ penalty enables to shrink less-informative coefficients $\beta$ to zero and thereby increases the simplicity and explanatory power of the model. Hyperparameter $\lambda$ in the objective function~(\ref{eq:log_reg_lasso_obj}) enables to set a trade-off between the quality of the fit and sparsity of the model.
Equation~(\ref{eq:log_reg}) is also used to derive predictions. Estimated values of regression coefficients $\hat{\beta}_0$ and $\hat{\beta}$ together with the observation $\textbf{x}$ are plugged into the right hand side of Eq.~(\ref{eq:log_reg}) and the resulting value is used as an estimate $\hat{y}$. Using hyperparameter $\theta$, the thresholding is used to transform $\hat{y}$ to the binary value. Hence, if $\hat{y} \geq \theta$ the prediction is $1$ and otherwise it is $0$.
\subsection{Random forests}
Random Forests (RF) is a method based on the regression tree model~\cite{Hastie_2009}. The regression tree model predicts the target variables from ramified observations. More specifically, this method splits the training data into several subsets by applying conditions upon predictors. The model training represents a ramification process. When the tree is trained, branches grow from a single node, and every node determines a condition on a single predictor. A unique path is traced on the basis of the value of a single predictor, iteratively splitting the datasets into two children subsets. In order to determine the local optimal condition for the split, the Mean Squared Error (MSE) is minimized. RF allows diversifying the training of multiple regression trees, \cite{Breiman2001}. Particularly, a number \textit{m} of individual trees (i.e. the RF) is independently trained using a bagged (bootstrap aggregated) subset of the total training data.
The \textit{m-}th regression tree generates a prediction through the following equation:
\begin{equation}
\hat{y}^{(m)} =\sum_{i=1}^{n} w_i^{(m)}(\textbf{x})\cdot y_i,
\label{eq:regrtree}
\end{equation}
where $\textbf{x}$ is an observation and ${w_i^{(m)}}(\textbf{x})$ is weight evaluated as follows:
\begin{equation}
w_i^{(m)}(\textbf{x})=\frac{1\{\textbf{x}_i\in R_{L^{(m)}}\}}{\sum_{j=1}^{n}1\{\textbf{x}_j\in R_{L^{(m)}}\}},
\label{eq:RFweights}
\end{equation}
with $L^{(m)}$ the leaf of the \textit{m-}th tree individuated by $\textbf{x}$ , and $R_{L^{(m)}}$ the domain of this leaf. The function $1(\cdot)$ takes value 1 if the expression within the brackets is true and 0 otherwise. Thus, the \textit{m}-th tree returns as a prediction the average value of all responses that belong to the same leaf node as the observation for which the prediction is made. Then, RF returns its prediction through a simple average of the predictions of the individual trees. Finally, to obtain binary prediction, thresholding is applied.
\subsection{Gradient boosting regression tree}
The Gradient boosting regression tree (GBRT) exploits regression trees, within a different framework. In this case, regression trees are fitted upon residuals of a weak learner in an iterative way. The fitting stops when the improvement brought by the last iteration is smaller than a fixed threshold~\cite{friedman2001}.
At the generic iteration \textit{m}, the prediction is given by a recursive equation:
\begin{equation}
F_m(\textbf{x})=F_{m-1}(\textbf{x})+\sum_{j=1}^{J^{(m)}}\gamma_j^{(m)} \cdot 1\{\textbf{x}\in R_{j^{(m)}}\},
\label{eq:GBoutcome}
\end{equation}
where \textit{F} is the prediction provided by the weak learner, $J^{(m)}$ is the number of terminal regions $R_1^{(m)}, \dots, R_{J^{(m)}}^{(m)}$ (each corresponding to one leaf node) and $\gamma_1^{(m)},\dots , \gamma_{J^{(m)}}^{(m)}$ are parameters estimated at each iteration \textit{m}. The estimation of these parameters is accomplished by minimizing the MSE. The equation (\ref{eq:GBoutcome}) gives the prediction of the target variable. Finally, to obtain binary prediction, thresholding is applied.
\section{Results}
\label{sec:results}
In this section, the predictability of the popularity of charging pools is evaluated using various metrics. Since the LR-$\mathit{l}_1$ classification method is returning a sparse vector of regression coefficients, we evaluate the type and strength of the influence of predictors on the popularity as well.
\subsection{Measures of predictability }
A set of measures (see Table~\ref{tab:measureEquations}) was compiled to assess the performance of classification models from different perspectives. All measures can be calculated from the elements TP (true positives), TN (true negatives), FP (false positives) and FN (false negatives) of the confusion matrix~\cite[p.254]{kuhn2013applied}.
\begin{table}[ht]
\centering
\setlength{\extrarowheight}{7pt}
\begin{tabular}{lc}
\hline
\makecell{Predictability \\ measure} & \makecell{Mathematical \\ expression}\\
\hline
\textit{Accuracy} & $\frac{TP + TN}{n}$ \\
\textit{Precision} & $\frac{TP}{TP + FP}$ \\
\textit{Sensitivity} & $\frac{TP}{TP + FN}$ \\
\textit{Fall--out} & $\frac{FP}{FP + TN}$ \\
\textit{F--score} & $2\frac{Sensitivity\cdot Precision}{Sensitivity + Precision}$ \\
\makecell{\textit{Matthews’ correlation} \\ \textit{coefficient} (MCC)}& $\frac{TP\cdot TN-FP\cdot FN}{\sqrt{(TP+FP)(TP+FN)(TN+FP)(TN+FN)}}$\\
\hline
\end{tabular}
\caption{Overview of predictability measures and mathematical descriptions showing how the predictability measures are calculated from the elements of the confusion matrix.}
\label{tab:measureEquations}
\end{table}
The \textit{accuracy} is the proportion of pools predicted correctly, the \textit{precision} is proportion of correct predictions of popular pools and the \textit{sensitivity} (also called true positive rate) is the proportion of popular pools predicted correctly. The \textit{fall--out} (called also false positive rate) is a fraction of unpopular pools predicted incorrectly and it is used on the x-axis in the receiver operating characteristic (ROC) curve. The overall performance of a classifier can be evaluated by the area under the ROC curve (AUC) \cite[p.~147]{james2013introduction}.
By analyzing the dependency between a measure of predictability and probabilistic threshold $\theta$ a suitable range for $\theta$ can be determined. When applying this approach to the accuracy for data with class imbalance (i.e. unequal number of $1$s and $0$s in the response vector) it may not be intuitive to determine models reaching good accuracy (e.g. for $25\%$ of $1$s in the response vectors, we reach value of accuracy $0.625$ by just randomly guessing $1$s with probability of $0.25$ and $0$s with probability of $0.75$). The F--score and Matthews’ correlation coefficient (MCC) are more balanced measures recommended for data with class imbalance~\cite{kuhn2013applied,matthews1975comparison}. The \textit{F--score} combines precision and sensitivity, taking harmonic mean of both measures, i.e. it moves towards the lower of the values~\cite{sasaki2007truth}. By definition, if any of the values in the parentheses in the denominator of the equation defining MCC is $0$ (see Table~\ref{tab:measureEquations}), the MCC is set to $0$. For models predicting better (worse) than a random model, the MCC is positive (negative). Models as skilled as a random guess have MCC equal to $0$. The MCC equals to $1$ ($-1$) if all the observations are predicted correctly (wrongly).
\subsection{Settings of parameters}
A radius of the buffer, representing vicinity of a pool, was chosen from the set $\{100, 150, 200, 250, 300, 350, 400, 450, 500\}$ meters. This range is the distance drivers typically walk from the parking place to their destination~\cite{Waerden_2015}. The dependency between the popularity of charging pools and predictors was fitted by the least squares model for all values in the set. The highest value of $R^2$ was found for the radius of $350$ meters and selected for further analyses. In experiments, training and testing sets are assigned $80\%$ and $20\%$, respectively, using stratified sampling with respect to the response variable~\cite[p.68]{kuhn2013applied}. To gain more reliable results and conclusions, we evaluate variability of calculated measures by analyzing outputs of 100 models, trained on 100 different splits into training and testing sets. While evaluating predictability measures, we varied threshold $\theta$ in the range from 0 to 0.99, in steps of 0.01. The hyperparameter $\lambda$ in Eq.~(\ref{eq:log_reg_lasso_obj}), was found by k-fold cross validation, where the stratified sampling was used to split data into $k=10$ folds preserving the class distribution of the response variable. Considering values $10^{-4 + i*0.015}$, for $i \in\{0, ..., 200\}$, $\lambda$ is assigned the value corresponding to the largest AUC value~\cite{friedman2010regularization}. Similarly, when growing decision trees, the k-fold cross validation was applied to set the number of learning cycles, the learn rate for shrinkage, the minimum size of leaves, and the maximum number of splits.
The values of predictability measures were obtained by applying the trained model to the testing dataset. The response vector was encoded into a binary format by setting $25\%$ of the elements corresponding to top charging pools ranked by the popularity to value $1$ and the rest to $0$. We did also experiments with values of $15\%$, $20\%$, $30\%$ and $35\%$, however, very similar conclusions could be drawn from the results and we do not report them. In computations, we used $\mathit{l}_1$-regularized logistic regression implemented by the R package \textit{glmnet}~\cite{friedman2010regularization}. GBRT and RF were implemented in MATLAB environment within Statistic and Machine Learning Toolbox.
\subsection{Predictability of popularity of charging pools}
\begin{figure}
\centering
\includegraphics[width=0.97\textwidth]{fig1}
\caption{The mean value of the accuracy, precision and sensitivity evaluated on testing data as a function of threshold $\theta$ in (A) for LR-$l_1$, in (B) for GBRT and in (C) for RF method. Each measure is displayed in different line style. Thick lines show average values over an ensemble of 100 different training and testing datasets splits and shaded areas represent one standard deviation. Thin grey lines show the expected value of measures, achieved by the random model predicting popular pools with probability of $0.25$.}
\label{fig:acc_prec_sens}
\end{figure}
\begin{figure}
\centering
\includegraphics[width=0.97\textwidth]{fig2}
\caption{Mean value of the F--score and MCC measures evaluated on testing data as a function of threshold $\theta$ in (A) for LR-$l_1$, in (B) for GBRT and in (C) for RF method. Lines show average values over an ensemble of 100 different training and testing datasets splits and shaded areas represent one standard deviation.Thin grey lines show the expected value of measures, achieved by the random model predicting popular pools with probability of $0.25$.}
\label{fig:MCC_YJS}
\end{figure}
In Figure~\ref{fig:acc_prec_sens}, the mean accuracy, precision and sensitivity are shown for all three methods. To facilitate evaluation of the quality of predictions, we consider a null model predicting popular stations randomly with probability of $0.25$. It can be easily shown that considered measures for the null model take the functional forms presented in Table~\ref{tab:nullModelEquations} and displayed in Figures~\ref{fig:acc_prec_sens} and~\ref{fig:MCC_YJS} by thin lines.
\begin{table}[ht]
\centering
\setlength{\extrarowheight}{7pt}
\begin{tabular}{lc}
\hline
Measure & Function \\
\hline
\textit{Accuracy} & $0.25 + 0.5 \theta $ \\
\textit{Precision} & $1 - \theta $ \\
\textit{Sensitivity} & $0.25$ \\
\textit{F--score} & $\frac{1-\theta}{2.5-2\theta}$ \\
\textit{MCC} & $0$ \\
\hline
\end{tabular}
\caption{Functional forms describing the expected value of evaluated measures for the considered null model as a function of the threshold~$\theta$.}
\label{tab:nullModelEquations}
\end{table}
All three methods outperform the null model, in accuracy and precision measures in the whole range of $\theta$. It is unavoidable that the sensitivity is decreasing as the threshold $\theta$ grows. For $\theta$ values greater than $0.5$ the sensitivity falls below 0.5, which results in low applicability of predictions. As $\theta$ grows, the sensitivity can reach value zero, even for an ideal model, if the threshold $\theta$ is already too high. If $\theta$ is larger than 0.5, for the decision tree methods we find the sensitivity to be smaller than the null model confirming that such thresholds are too high.
In the literature, various approaches on how to select the value of the threshold $\theta$ and hence to find a reasonable trade-off between accuracy, precision and sensitivity can be found~\cite{kuhn2013applied}. We selected two metrics, MCC and F--score, which are evaluated in Figure~\ref{fig:MCC_YJS}. Both measures reach one single maximum, which is hence also the global maximum. Values of the threshold where the maxima are achieved we denote as $\theta_{MCC_{max}}$ and $\theta_{F\textendash score_{max}}$. Values of the accuracy, precision and sensitivity for $\theta_{MCC_{max}}$ and $\theta_{F\textendash score_{max}}$ are reported in Table~\ref{tab:thetaResults}.
\begin{table}[ht]
\centering
\begin{tabular}{r|rrr}
\hline
& LR-$l_1$ & GBRT & RF \\
\hline
$\theta_{MCC_{max}}$ & 0.34 & 0.35 & 0.37 \\
\hline
Accuracy & 0.829 & 0.823 & 0.819 \\
Precision & 0.648 & 0.644 & 0.632 \\
Sensitivity & 0.728 & 0.694 & 0.701 \\
MCC & 0.571 & 0.548 & 0.542 \\
\hline
$\theta_{F\textendash score_{max}}$ & 0.34 & 0.35 & 0.35 \\
\hline
Accuracy & 0.829 & 0.813 & 0.813 \\
Precision & 0.633 & 0.614 & 0.615 \\
Sensitivity & 0.750 & 0.736 & 0.727 \\
F--score & 0.685 & 0.667 & 0.664 \\
\hline
\end{tabular}
\caption{The mean values of the accuracy, precision and sensitivity from Figure~\ref{fig:acc_prec_sens} for all three methods and selected threshold values $\theta_{MCC_{max}}$ (threshold value corresponding to the maximum value of the MCC measure) and $\theta_{F\textendash score_{max}}$ (threshold value corresponding to the maximum value of the $F\textendash score$ measure).}
\label{tab:thetaResults}
\end{table}
When decisions about the extension of the existing charging infrastructure are taken, many often contradicting factors are taken into account. Alternatively, stakeholders can select a threshold $\theta$ based on their expectations and attitudes towards risk. The lower the $\theta$, the more likely is identification of popular locations with the drawback of increased risk of placing charging pools into unpopular areas. In opposite, the higher $\theta$, we identify the popular charging pools with higher assurance, with the drawback of overlooking potentially popular locations.
According to observed values of measures, we recommend considering the threshold $\theta$ within the range from 0.3 to 0.45, where both precision and sensitivity are relatively high.
Comparison with the null model and values of measures in Table~\ref{tab:thetaResults} indicate that urban area and characteristics of charging pools contain some predictive power for the popularity. What values of measure justify good quality of models typically depend on the application domain~\cite[p.70]{james2013introduction}. Considering data analyses concerning the human choice in similar domains such as for example bike sharing applications~\cite{zhang2016bicycle, chen2017understanding}, the obtained values of the accuracy exceeding value 0.8, while both precision and sensitivity are larger than 0.65, can be considered as favorable.
\begin{figure}
\centering
\includegraphics[width=0.97\textwidth]{fig3}
\caption{Ensemble of 100 ROC curves (sensitivity versus fall-out) corresponding to different training and testing datasets splits. The dashed line visualizes points corresponding to the random classifier. Value of the area under the curve (AUC) and the corresponding standard deviation are reported in the frame located at the bottom.}
\label{fig:ROC}
\end{figure}
Simple inspection of Figures~\ref{fig:acc_prec_sens}-~\ref{fig:MCC_YJS} indicates, that all three methods provide similar results. To evaluate, whether the results are statistically distinguishable, we test the differences in AUC ensembles, i.e. areas under the ROC curves. The ROC curves corresponding to 100 different training and testing dataset splits are shown in Figure~\ref{fig:ROC}. In all statistical tests we set the significance level $\alpha = 0.01$.
In the first step, the equality of AUC variances between all pairs of methods was tested using the F-test. Statistically significant difference was identified only for LR-$\mathit{l}_1$ and RF methods, the first having less variance ($p = 0.0017$). To compare mean values of AUCs, for the case when the variances were not found to be significantly different, the t-test was used otherwise, the Welch t-test \cite{welch1947generalization}. The mean AUC of LR-$\mathit{l}_1$ was larger than RF ($p = 0.0094$), and not different from GBRT ($p = 0.0379$). The mean AUCs of RF and GBRT were not significantly different ($p = 0.5091$). Hence, according to AUC values, the method LR-$\mathit{l}_1$ is more stable than RF, and outperforms the method in mean values of AUC. On the significance level $\alpha = 0.05$, LR-$\mathit{l}_1$ method has a significantly larger mean value of AUCs than both, the GBRT and RF methods.
\subsection{Characteristics influencing the popularity of public charging stations}
The shrinkage nature of $\mathit{l}_1$ penalty in the LR-$\mathit{l}_1$ method provides better predictions on testing data and variable selection functionality. This gives an advantage compared to decision tree methods, that yielded clumsy models, involving 169 predictors on average, making them hard to interpret. Due to this reason, we interpret the results for the LR-$\mathit{l}_1$ method only.
The values $\hat{\beta}$ of the LR-$\mathit{l}_1$ model might be sensitive to the used sample of training data, hence an analysis of the results robustness is necessary. The conventional solution is to evaluate the $p$-values of coefficients estimated by statistical methods. The problem to calculate $p$-values for LR-$\mathit{l}_1$ model is difficult due to the adaptive nature of the estimation procedure~\cite{tibshirani2015statistical}. Therefore, to capture the stochasticity in coefficients $\hat{\beta}$, we estimate their distributions by sampling the dataset 500 times using bootstrap~\cite[p.~187]{james2013introduction}. A model is fitted to each bootstrapped dataset using stratified cross-validation. To facilitate a comparison of the impact of predictors, we standardize each element of $\hat{\beta}$ by dividing it with the sample standard deviation of the corresponding predictor.
The group of selected predictors depends on the training data sample. In Figure~\ref{fig:all_var_importance}, sampled distributions of coefficients that have been selected most frequently (at least by 90\% out of 500 models) are displayed. The bar plot presents the percentage of models where the coefficients were equal to zero.
\begin{figure}
\centering
\includegraphics[width=0.97\textwidth]{fig4}
\caption{
(A) Tukey Box-plot of standardized coefficients $\hat{\beta_j}$ for 500 $LR-l1$ models. The coefficients are sorted by median values. Only predictors that were selected by at least 90\% of models are displayed. (B) The percentage of models setting the standardized coefficient $\hat{\beta_j}$ to zero.
}
\label{fig:all_var_importance}
\end{figure}
The impact of a predictor on the response increases with the absolute value of the corresponding coefficient~\cite{kuhn2013applied}. Positive (negative) sign of a coefficient indicates increasing (decreasing) impact of the predictor. A way how to quantify the significance of coefficient $\hat{\beta}$ is to assess the likelihood that the coefficient is different from zero. Analysis of the distribution of coefficients $\hat{\beta}$ allows us to make such assessments and to conclude how certain is the positive or negative influence of predictors on the response variable.
In summary, selected predictors can be categorized into three groups: the function of the geographic area constituting the vicinity of a charging pool, characteristics of the population living in this area and properties of charging pools. From the perspective of the geographic area, the most important predictors are the number of wholesale businesses, shops, hotels, restaurants and catering businesses, areas with recreational inland water, sports fields and roads, all having a positive impact. The minimal distance to financial, cultural, and transportation OpenStreetMap (OSM) amenities have negative coefficients, meaning that large minimal distance is decreasing the popularity of charging pools. Hence, if these facilities are found in the proximity of charging pools, they have a positive impact on the popularity.
In opposite, the residential areas, the areas with non-commercial ornamental and vegetable cultivation and the presence of OSM amenities related to households tend to lower the popularity of charging pools. These findings are well aligned with the intuition that in residential areas the charging pools are visited by more homogeneous groups of users than in the more crowded urban areas. Similarly, the most likely explanation of the negative impact of business and industrial areas is work charging~\cite{Sadeghianpourhamami_2018}, i.e. either charging of a fleet of company cars or regular use of charging pools by (a small group of) employees commuting to work.
A notable population group living near popular charging pools is working elderly people between 65 and 74 years old. In opposite, popularity is negatively correlated with areas inhabited by the population working in the mining, manufacturing or construction sector as well as persons who depend on social assistance. Thus, these results suggest that the economic prosperity of the population in the vicinity is affecting the visiting patterns of charging pools. The popular charging pools are more likely to be deployed following strategical rollout and have larger maximal power and more connectors. The negative influence of geographic longitude can be explained by the geography of the Netherlands, while the western part of the country is more urbanized and we can find here the majority of large Dutch cities.
\section{Discussion and conclusions}
\label{sec:discussion_and_conslusions}
This study demonstrates the ability of classification methods to predict popular locations of charging pools from various GIS and large-scale charging infrastructure data. Moreover, we evaluate the impact and significance of factors affecting the popularity of charging infrastructure. Predicting the popularity of charging pools is of utmost importance for matching EV requirements driven by social habits with energy requirements related to electrical network configuration. Main conclusions derived from the data analysis are the following:
\begin{itemize}
\item
To characterize the performance of charging infrastructure is a complex task as various viewpoints need to be considered. We selected seven indicators characterizing performance considering energy supply issues, charging demand and expectations of infrastructure operators. The popularity of charging pools represented by the number of unique RFID cards can be explained to the largest degree among the seven indicators of charging infrastructure.
\item To exploit the previous finding, we formulate the classification problem of determining top charging pools, ranked by the popularity. The $\mathit{l}_1$-regularized logistic regression, gradient boosted decision trees and random forests, are able to predict the popularity with the accuracy exceeding value 0.8 and F--score reaching value 0.68, clearly outperforming random models. Such values do not justify decisions taken solely based on predictions provided by created models, however, results from our models can give indications for decision making processes in charging infrastructure planning.
\item Factors having positive influence on the popularity of charging pools are rollout strategy, maximal power, number of connectors, number of shopping, catering, sport and trade related venues. Hence, charging pools located on frequently visited spots and providing convenient charging opportunities are more likely to become popular. The largest negative influence have residential, business and industrial areas and areas inhabited by workers and persons receiving social benefits. Thus, areas inhabited by social groups with lower purchase power and areas periodically visited by a small group of EV drivers are associated with lower popularity.
\end{itemize}
The presented results are limited by the low utilization of some charging pools, which can be attributed to the low penetration of electric vehicles in some areas in the Netherlands. Thanks to the rapid growth of EVs, the utilization of charging pools is expected to increase and potentially mitigate this limitation. A difficulty often present in GIS data is the interdependence of factors, expressed as collinearity, causing nontrivial problems when interpreting the impact of individual factors. There is no generally accepted approach to address this problem. We minimize the chances that collinearity affects the results by removing highly correlated factors. Nevertheless, predictors that we present as influential and significant, should be taken into account with some care. Typically, the uncaptured stochasticity of models can be attributed to missing data. We assume that more detailed mobility data, e.g. GPS and floating car data could improve the results. Geographically, our study is focused on the area of the Netherlands, which might impose some limitations when transferring models and conclusions to other countries. Although, we expect similar results for comparable geographic and demographic contexts.
This study opens several directions. For instance, future research could explore possibilities how to design prediction models for other performance indicators of charging pools, how to efficiently downscale prediction models to a level of a region or a city or how to improve predictions by customizing models to specific classes of charging pools. Another challenge is the application of regression approaches that could successfully predict the values of performance indicators.
\section*{Acknowledgements}
This work was supported by the research grants: VEGA 1/0089/19 ”Data analysis methods and decisions support tools for service systems supporting electric vehicles”, VEGA 1/342/18 ”Optimal dimensioning of service systems”, APVV-15-0179 ”Reliability of emergency systems on infrastructure with uncertain functionality of critical elements”, Operational Program Research and Innovation in frame of the project: ICTproducts for intelligent systems communication, code ITMS2014+ 313011T413, co-financed by the European Regional Development Fund and by the Slovak Research and Development Agency under the contract no. SK-IL-RD-18-005. We thank \v{L}udmila J\'{a}no\v{s}\'{i}kov\'{a} and Luca Lena Jansen for valuable comments and suggestions.
\bibliographystyle{cas-model2-names}
\begin{thebibliography}{65}
\expandafter\ifx\csname natexlab\endcsname\relax\def\natexlab#1{#1}\fi
\ifx\xfnm\relax \def\xfnm[#1]{\unskip,\space#1}\fi
\bibitem[{Ajanovic and Haas(2016)}]{ajanovic2016dissemination}
Ajanovic, A., Haas, R.,
2016.
\newblock Dissemination of electric vehicles in urban areas:
Major factors for success.
\newblock Energy 115,
1451--1458.
\newblock doi:10.1016/j.energy.2016.05.040.
\bibitem[{Asamer et~al.(2016)Asamer, Reinthaler, Ruthmair, Straub and
Puchinger}]{Asamer2016}
Asamer, J., Reinthaler, M.,
Ruthmair, M., Straub, M.,
Puchinger, J., 2016.
\newblock Optimizing charging station locations for urban taxi
providers.
\newblock Transportation Research Part A: Policy and
Practice 85, 233–246.
\newblock doi:10.1016/j.tra.2016.01.014.
\bibitem[{Boyd and Vandenberghe(2004)}]{Boyd_2004}
Boyd, S., Vandenberghe, L.,
2004.
\newblock Convex Optimization.
\newblock Cambridge University Press,
New York.
\bibitem[{Breiman(2001)}]{Breiman2001}
Breiman, L., 2001.
\newblock Random forests.
\newblock Machine Learning 45,
5–32.
\bibitem[{Buzna et~al.(2019)Buzna, {De Falco}, Khormali, Proto and
Straka}]{Buzna2019}
Buzna, L., {De Falco}, P.,
Khormali, S., Proto, D.,
Straka, M., 2019.
\newblock Electric vehicle load forecasting: A comparison
between time series and machine learning approaches, in:
2019 1st International Conference on Energy Transition in
the Mediterranean Area (SyNERGY MED), IEEE. p.
1–5.
\bibitem[{{CBS Netherlands}(2015a)}]{lccbs}
{CBS Netherlands}, 2015a.
\newblock Cbs land cover.
\newblock
\texttt{https://www.pdok.nl/downloads?articleid=1951731}.
\newblock Accessed: 2016-09-14.
\bibitem[{{CBS Netherlands}(2015b)}]{neighbpop}
{CBS Netherlands}, 2015b.
\newblock Neighbourhoods dataset 2015.
\newblock
\texttt{https://www.cbs.nl/nl-nl/dossier/nederland-regionaal/geografische\%20data/wijk-en-buurtkaart-2015}.
\newblock Accessed: 2018-08-20.
\bibitem[{{CBS Netherlands}(2015c)}]{popcores}
{CBS Netherlands}, 2015c.
\newblock Population cores in the netherlands.
\newblock
\texttt{https://www.cbs.nl/nl-nl/achtergrond/2014/13/bevolkingskernen-in-nederland-2011}.
\newblock Accessed: 2017-09-14.
\bibitem[{Chen et~al.(2017)Chen, Ma, Pan, Jakubowicz
et~al.}]{chen2017understanding}
Chen, L., Ma, X., Pan,
G., Jakubowicz, J., et~al., 2017.
\newblock Understanding bike trip patterns leveraging bike
sharing system open data.
\newblock Frontiers of computer science
11, 38–48.
\bibitem[{Chen et~al.(2015)Chen, Zhang, Pan, Ma, Yang, Kushlev, Zhang and
Li}]{Chen2015}
Chen, L., Zhang, D., Pan,
G., Ma, X., Yang, D.,
Kushlev, K., Zhang, W.,
Li, S., 2015.
\newblock Bike sharing station placement leveraging
heterogeneous urban open data, in: Proceedings of the
2015 ACM International Joint Conference on Pervasive and Ubiquitous
Computing, ACM. p. 571–575.
\bibitem[{Chen et~al.(2013)Chen, Kockelman and Khan}]{Chen2013}
Chen, T.D., Kockelman, K.M.,
Khan, M., 2013.
\newblock Locating electric vehicle charging stations:
Parking-based assignment method for seattle, washington.
\newblock Transportation Research Record
2385, 28–36.
\bibitem[{Coffman et~al.(2017)Coffman, Bernstein and Wee}]{Coffman2017}
Coffman, M., Bernstein, P.,
Wee, S., 2017.
\newblock Electric vehicles revisited: a review of factors that
affect adoption.
\newblock Transport Reviews 37,
79–93.
\bibitem[{Davidov and Pantoš(2017)}]{Davidov2017}
Davidov, S., Pantoš, M.,
2017.
\newblock Planning of electric vehicle infrastructure based on
charging reliability and quality of service.
\newblock Energy 118,
1156–1167.
\bibitem[{{De Gennaro} et~al.(2015){De Gennaro}, Paffumi and Martini}]{De2015}
{De Gennaro}, M., Paffumi, E.,
Martini, G., 2015.
\newblock Customer-driven design of the recharge infrastructure
and vehicle-to-grid in urban areas: A large-scale application for electric
vehicles deployment.
\newblock Energy 82,
294–311.
\bibitem[{Develder et~al.(2016)Develder, Sadeghianpourhamami, Strobbe and
Refa}]{Develder2016}
Develder, C., Sadeghianpourhamami, N.,
Strobbe, M., Refa, N.,
2016.
\newblock Quantifying flexibility in ev charging as dr
potential: Analysis of two real-world data sets, in:
2016 IEEE International Conference on Smart Grid
Communications (SmartGridComm), IEEE. p.
600–605.
\bibitem[{Dong et~al.(2014)Dong, Liu and Lin}]{Dong2014}
Dong, J., Liu, C., Lin,
Z., 2014.
\newblock Charging infrastructure planning for promoting
battery electric vehicles: An activity-based approach using multiday travel
data.
\newblock Transportation Research Part C: Emerging
Technologies 38, 44–55.
\bibitem[{D'Silva et~al.(2018)D'Silva, Noulas, Musolesi, Mascolo and
Sklar}]{DSilva2018}
D'Silva, K., Noulas, A.,
Musolesi, M., Mascolo, C.,
Sklar, M., 2018.
\newblock Predicting the temporal activity patterns of new
venues.
\newblock EPJ Data Science 7,
13.
\newblock URL: \texttt{https://doi.org/10.1140/epjds/s13688-018-0142-z},
doi:10.1140/epjds/s13688-018-0142-z.
\bibitem[{EC(2019)}]{JRC_web}
EC, 2019.
\newblock 2030 climate \& energy framework.
\newblock
https://ec.europa.eu/clima/policies/strategies/2030\_en.
\bibitem[{Efthymiou et~al.(2017)Efthymiou, Chrysostomou, Morfoulaki and
Aifantopoulou}]{Efthymiou2017}
Efthymiou, D., Chrysostomou, K.,
Morfoulaki, M., Aifantopoulou, G.,
2017.
\newblock Electric vehicles charging infrastructure location: a
genetic algorithm approach.
\newblock European Transport Research Review
9, 27.
\bibitem[{(EIA).(2013)}]{EIA2013}
(EIA)., U.E.I.A., 2013.
\newblock International energy outlook 2016 with projections to
2040.
\bibitem[{{ElaadNL}(2015)}]{elaadnl}
{ElaadNL}, 2015.
\newblock Elaadnl.
\newblock \texttt{https://www.elaad.nl/}.
\newblock Accessed: 2019-06-21.
\bibitem[{EnergieAtlas(2015)}]{energyatlas}
EnergieAtlas, N., 2015.
\newblock Energy atlas.
\newblock
\texttt{https://www.pdok.nl/downloads?articleid=1951681}.
\newblock Accessed: 2018-10-16.
\bibitem[{Flammini et~al.(2017)Flammini, Prettico, Fulli, Bompard and
Chicco}]{Flammini2017}
Flammini, M.G., Prettico, G.,
Fulli, G., Bompard, E.,
Chicco, G., 2017.
\newblock Interaction of consumers, photovoltaic systems and
electric vehicle energy demand in a reference network model, in:
2017 International Conference of Electrical and
Electronic Technologies for Automotive, IEEE. p.
1–5.
\bibitem[{Flammini et~al.(2019)Flammini, Prettico, Julea, Fulli, Mazza and
Chicco}]{Flammini2019}
Flammini, M.G., Prettico, G.,
Julea, A., Fulli, G.,
Mazza, A., Chicco, G.,
2019.
\newblock Statistical characterisation of the real transaction
data gathered from electric vehicle charging stations.
\newblock Electric Power Systems Research
166, 136–150.
\bibitem[{Friedman(2001)}]{friedman2001}
Friedman, J., 2001.
\newblock Greedy function approximation: a gradient boosting
machine.
\newblock Annals of statistics 29,
1189–1232.
\bibitem[{Friedman et~al.(2010)Friedman, Hastie and
Tibshirani}]{friedman2010regularization}
Friedman, J., Hastie, T.,
Tibshirani, R., 2010.
\newblock Regularization paths for generalized linear models
via coordinate descent.
\newblock Journal of statistical software
33, 1--22.
\bibitem[{Ge et~al.(2012)Ge, Liang, Hong and Long}]{GE2012}
Ge, S.y., Liang, F., Hong,
L., Long, W., 2012.
\newblock The planning of electric vehicle charging stations in
the urban area, in: 2nd International Conference on
Electronic \& Mechanical Engineering and Information Technology,
Atlantis Press. pp. 1598 -- 1604.
\bibitem[{Gonzalez et~al.(2014)Gonzalez, Alvaro, Gamallo, Fuentes,
Fraile-Ardanuy, Knapen and Janssens}]{Gonzalez2014}
Gonzalez, J., Alvaro, R.,
Gamallo, C., Fuentes, M.,
Fraile-Ardanuy, J., Knapen, L.,
Janssens, D., 2014.
\newblock Determining electric vehicle charging point locations
considering drivers’ daily activities.
\newblock Procedia Computer Science 32,
647–654.
\bibitem[{Hastie et~al.(2009)Hastie, Tibshirani and Friedman}]{Hastie_2009}
Hastie, T., Tibshirani, R.,
Friedman, J., 2009.
\newblock The elements of statistical learning: data mining,
inference and prediction.
\newblock 2 ed., Springer.
\newblock URL: \texttt{http://www-stat.stanford.edu/~tibs/ElemStatLearn/}.
\bibitem[{Hastie et~al.(2015)Hastie, Tibshirani and Wainwright}]{Hastie_2015}
Hastie, T., Tibshirani, R.,
Wainwright, M., 2015.
\newblock Statistical learning with sparsity: The lasso and
generalizations.
\newblock CRC Press.
\newblock doi:10.1201/b18401.
\bibitem[{He et~al.(2018)He, Yang, Tang and Huang}]{He2018}
He, J., Yang, H., Tang,
T.Q., Huang, H.J., 2018.
\newblock An optimal charging station location model with the
consideration of electric vehicle’s driving range.
\newblock Transportation Research Part C: Emerging
Technologies 86, 641–654.
\bibitem[{Helmus et~al.(2018)Helmus, Spoelstra, Refa, Lees and van~den
Hoed}]{Helmus2018}
Helmus, J., Spoelstra, J.,
Refa, N., Lees, M.,
van~den Hoed, R., 2018.
\newblock Assessment of public charging infrastructure push and
pull rollout strategies: The case of the netherlands.
\newblock Energy policy 121,
35–47.
\bibitem[{Van~den Hoed et~al.(2013)Van~den Hoed, Helmus, De~Vries and
Bardok}]{van2013data}
Van~den Hoed, R., Helmus, J.,
De~Vries, R., Bardok, D.,
2013.
\newblock Data analysis on the public charge infrastructure in
the city of amsterdam.
\newblock World Electric Vehicle Journal
6, 829--838.
\newblock doi:10.3390/wevj6040829.
\bibitem[{James et~al.(2013)James, Witten, Hastie and
Tibshirani}]{james2013introduction}
James, G., Witten, D.,
Hastie, T., Tibshirani, R.,
2013.
\newblock An introduction to statistical learning. volume
112.
\newblock Springer.
\bibitem[{Karamshuk et~al.(2013)Karamshuk, Noulas, Scellato, Nicosia and
Mascolo}]{Karamshuk2013}
Karamshuk, D., Noulas, A.,
Scellato, S., Nicosia, V.,
Mascolo, C., 2013.
\newblock Geo-spotting: mining online location-based services
for optimal retail store placement, in: Proceedings of
the 19th ACM SIGKDD international conference on Knowledge discovery and data
mining, ACM. p. 793–801.
\bibitem[{Kuhn and Johnson(2013)}]{kuhn2013applied}
Kuhn, M., Johnson, K.,
2013.
\newblock Applied predictive modeling.
volume~26.
\newblock Springer.
\bibitem[{Liu et~al.(2012)Liu, Wen and Ledwich}]{Liu2012}
Liu, Z., Wen, F., Ledwich,
G., 2012.
\newblock Optimal planning of electric-vehicle charging
stations in distribution systems.
\newblock IEEE Transactions on Power Delivery
28, 102–110.
\bibitem[{Lucas et~al.(2019)Lucas, Barranco and Refa}]{Lucas2019}
Lucas, A., Barranco, R.,
Refa, N., 2019.
\newblock Ev idle time estimation on charging infrastructure,
comparing supervised machine learning regressions.
\newblock Energies 12,
269.
\bibitem[{Lucas et~al.(2018)Lucas, Prettico, Flammini, Kotsakis, Fulli and
Masera}]{Lucas2018}
Lucas, A., Prettico, G.,
Flammini, M.G., Kotsakis, E.,
Fulli, G., Masera, M.,
2018.
\newblock Indicator-based methodology for assessing ev charging
infrastructure using exploratory data analysis.
\newblock Energies 11.
\newblock URL: \texttt{http://www.mdpi.com/1996-1073/11/7/1869},
doi:10.3390/en11071869.
\bibitem[{Matthews(1975)}]{matthews1975comparison}
Matthews, B.W., 1975.
\newblock Comparison of the predicted and observed secondary
structure of t4 phage lysozyme.
\newblock Biochimica et Biophysica Acta (BBA)-Protein
Structure 405, 442–451.
\bibitem[{{Ministry of the Interior and Kingdom Relations}(2015)}]{liveability}
{Ministry of the Interior and Kingdom Relations},
2015.
\newblock Liveability meter.
\newblock
\texttt{https://data.overheid.nl/data/dataset/leefbaarometer-2-0---meting-2016}.
\newblock Accessed: 2018-10-15.
\bibitem[{{Netherlands Enterprise Agency}(2019)}]{EVdefinitions}
{Netherlands Enterprise Agency}, 2019.
\newblock Electric vehicle charging - definitions and
explanation.
\newblock
\texttt{https://www.nklnederland.nl/uploads/files/Electric_Vehicle_Charging_-_Definitions_and_Explanation_-_january_2019.pdf}.
\newblock Accessed: 2019-06-19.
\bibitem[{{Oak Ridge National Laboratory}(2015)}]{landscan}
{Oak Ridge National Laboratory}, 2015.
\newblock Landscan datasets.
\newblock
\texttt{https://landscan.ornl.gov/landscan-datasets}.
\newblock Accessed: 2015-05-20.
\bibitem[{{OpenChargeMap}(2015)}]{ocm}
{OpenChargeMap}, 2015.
\newblock \texttt{https://openchargemap.org}.
\newblock Accessed: 2019-01-10.
\bibitem[{{OpenStreetMap}(2015)}]{osm}
{OpenStreetMap}, 2015.
\newblock \texttt{https://www.openstreetmap.org}.
\newblock Accessed: 2019-02-13.
\bibitem[{{Oplaadpalen}(2015)}]{opp}
{Oplaadpalen}, 2015.
\newblock \texttt{https://www.oplaadpalen.nl/}.
\newblock Accessed: 2019-02-20.
\bibitem[{Pevec et~al.(2018)Pevec, Babic, Kayser, Carvalho, Ghiassi-Farrokhfal
and Podobnik}]{Pevec2018}
Pevec, D., Babic, J.,
Kayser, M.A., Carvalho, A.,
Ghiassi-Farrokhfal, Y., Podobnik, V.,
2018.
\newblock A data-driven statistical approach for extending
electric vehicle charging infrastructure.
\newblock International journal of energy research
42, 3102–3120.
\bibitem[{Rahman et~al.(2016)Rahman, Vasant, Singh, Abdullah-Al-Wadud and
Adnan}]{Rahman2016}
Rahman, I., Vasant, P.M.,
Singh, B.S.M., Abdullah-Al-Wadud, M.,
Adnan, N., 2016.
\newblock Review of recent trends in optimization techniques
for plug-in hybrid, and electric vehicle charging infrastructures.
\newblock Renewable and Sustainable Energy Reviews
58, 1039–1047.
\bibitem[{Ren et~al.(2019)Ren, Zhang, Hu and Qiu}]{ren2019location}
Ren, X., Zhang, H., Hu,
R., Qiu, Y., 2019.
\newblock Location of electric vehicle charging stations: A
perspective using the grey decision-making model.
\newblock Energy 173,
548--553.
\bibitem[{Sadeghi-Barzani et~al.(2014)Sadeghi-Barzani, Rajabi-Ghahnavieh and
Kazemi-Karegar}]{Sadeghi2014}
Sadeghi-Barzani, P., Rajabi-Ghahnavieh, A.,
Kazemi-Karegar, H., 2014.
\newblock Optimal fast charging station placing and sizing.
\newblock Applied Energy 125,
289–299.
\bibitem[{Sadeghianpourhamami et~al.(2018a)Sadeghianpourhamami, Refa, Strobbe
and Develder}]{Sadeghianpourhamami2018}
Sadeghianpourhamami, N., Refa, N.,
Strobbe, M., Develder, C.,
2018a.
\newblock Quantitive analysis of electric vehicle flexibility:
A data-driven approach.
\newblock International Journal of Electrical Power \& Energy
Systems 95, 451–462.
\bibitem[{Sadeghianpourhamami et~al.(2018b)Sadeghianpourhamami, Refa, Strobbe
and Develder}]{Sadeghianpourhamami_2018}
Sadeghianpourhamami, N., Refa, N.,
Strobbe, M., Develder, C.,
2018b.
\newblock Quantitive analysis of electric vehicle flexibility:
A data-driven approach.
\newblock International Journal of Electrical Power \& Energy
Systems 95, 451–462.
\newblock URL:
\texttt{http://www.sciencedirect.com/science/article/pii/S0142061516323687},
doi:10.1016/j.ijepes.2017.09.007.
\bibitem[{Sasaki et~al.(2007)}]{sasaki2007truth}
Sasaki, Y., et~al., 2007.
\newblock The truth of the f-measure .
\bibitem[{Tao et~al.(2018)Tao, Huang and Yang}]{tao2018data}
Tao, Y., Huang, M., Yang,
L., 2018.
\newblock Data-driven optimized layout of battery electric
vehicle charging infrastructure.
\newblock Energy 150,
735--744.
\bibitem[{Tibshirani({1996})}]{Tibshirani_1996}
Tibshirani, R., {1996}.
\newblock Regression shrinkage and selection via the lasso.
\newblock {JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES
B-METHODOLOGICAL} {58}, {267–288}.
\bibitem[{Tibshirani et~al.(2015)Tibshirani, Wainwright and
Hastie}]{tibshirani2015statistical}
Tibshirani, R., Wainwright, M.,
Hastie, T., 2015.
\newblock Statistical learning with sparsity: the lasso and
generalizations.
\newblock Chapman and Hall/CRC.
\bibitem[{{Traffic flows}(2015)}]{traffic_flows}
{Traffic flows}, 2015.
\newblock The database was provided for reserach purposes by
the national institute for public healt and environment.
\newblock \texttt{http://www.rivm.nl/}.
\newblock Accessed: 2019-01-07.
\bibitem[{Trigg and Telleen(2013)}]{Tali_2013}
Trigg, T., Telleen, P.,
2013.
\newblock GLOBAL EV OUTLOOK: Understanding the Electric Vehicle
Landscape to 2020.
\newblock Technical Report. International Energy Agency.
\bibitem[{Verma et~al.(2015)Verma, Asadi, Yang and Tyagi}]{Verma2015}
Verma, A., Asadi, A.,
Yang, K., Tyagi, S.,
2015.
\newblock A data-driven approach to identify households with
plug-in electrical vehicles (pevs).
\newblock Applied energy 160,
71–79.
\bibitem[{van~der Waerden et~al.(2015)van~der Waerden, Timmermans and
de~Bruin-Verhoeven}]{Waerden_2015}
van~der Waerden, P., Timmermans, H.,
de~Bruin-Verhoeven, M., 2015.
\newblock Car drivers’ characteristics and the maximum
walking distance between parking facility and final destination.
\newblock Journal of Transport and Land Use
10, 1–11.
\newblock URL:
\texttt{https://www.jtlu.org/index.php/jtlu/article/view/568},
doi:10.5198/jtlu.2015.568.
\bibitem[{Welch(1947)}]{welch1947generalization}
Welch, B.L., 1947.
\newblock The generalization ofstudent's' problem when several
different population variances are involved.
\newblock Biometrika 34,
28–35.
\newblock doi:10.2307/2332510.
\bibitem[{Wilbanks and Greene(2010)}]{wilbanks2010}
Wilbanks, T.J., Greene, D.L.,
2010.
\newblock The importance of advancing technology to america's
energy goals.
\newblock Energy Policy 38.
\bibitem[{Xydas et~al.(2016)Xydas, Marmaras, Cipcigan, Jenkins, Carroll and
Barker}]{Xydas2016}
Xydas, E., Marmaras, C.,
Cipcigan, L.M., Jenkins, N.,
Carroll, S., Barker, M.,
2016.
\newblock A data-driven approach for characterising the
charging demand of electric vehicles: A uk case study.
\newblock Applied energy 162,
763–771.
\bibitem[{Yang et~al.(2017)Yang, Dong and Hu}]{Yang2017}
Yang, J., Dong, J., Hu,
L., 2017.
\newblock A data-driven optimization-based approach for siting
and sizing of electric taxi charging stations.
\newblock Transportation Research Part C: Emerging
Technologies 77, 462–477.
\bibitem[{Zhang et~al.(2016)Zhang, Pan, Li and Philip}]{zhang2016bicycle}
Zhang, J., Pan, X., Li,
M., Philip, S.Y., 2016.
\newblock Bicycle-sharing system analysis and trip prediction,
in: 2016 17th IEEE international conference on mobile
data management (MDM), IEEE. p.
174–179.
\end{thebibliography}
\section*{Author contributions statement}
CRediT (Contributor Roles Taxonomy) has been applied to describe contributions of authors. \textbf{Milan Straka:} Data curation, Methodology, Formal analysis, Investigation, Validation, Visualization, Software, Writing-original draft; \textbf{Pasquale De Falco:} Formal analysis, Methodology, Investigation, Software, Validation, Writing-original draft; \textbf{Gabriela Ferruzzi:} Validation; Writing original draft; Writing-review and editing; \textbf{Daniela Proto:} Validation; Writing original draft; Writing-review and editing; \textbf{Gijs van der Poel:} Conceptualization, Resources, Writing-review and editing; \textbf{Shahab Khormali:} Resources, Writing-original draft, Writing-review and editing;
\textbf{\v{L}ubo\v{s} Buzna:} Conceptualization, Funding acquisition, Methodology, Resources, Supervision, Writing-review and editing.
\section*{Additional information}
Competing financial interests: The authors declare no competing financial interests.
\newpage
\setcounter{section}{0}