Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,396 characters · 34 sections · 71 citation commands
Explaining the distribution of energy consumption at slow charging infrastructure for electric vehicles from socio-economic data
\let\WriteBookmarks\relax
\shorttitle{Explaining the distribution of energy consumption at slow charging infrastructure for electric vehicles from socio-economic data} \shortauthors{Straka, M. et al.}
\tnotemark[1,2] \tnotetext[1]{This document is a result of the research project VEGA 1/0089/19 “Data analysis methods and decisions support tools for service systems supporting electric vehicles”. It was partly supported by VEGA 1/342/18 “Optimal dimensioning of service systems”, APVV-19-0441 “Allocation of limited resources to public service systems with conflicting quality criteria”, Operational Program Research and Innovation in the frame of the project: ICT products for intelligent systems communication, code ITMS2014+ 313011T413, co-financed by the European Regional Development Fund and by the Slovak Research and Development Agency under the contract no. SK-IL-RD-18-005.}
[type=editor, auid=000,bioid=1, role=Researcher, orcid=0000-0002-8558-0408] \fnmark[1] \ead{[email removed]} \credit{Data curation, Methodology, Formal analysis, Investigation, Software, Validation, Visualization, Writing -- original draft} \address[1]{Department of Mathematical Methods and Operations Research, University of \v{Z}ilina, Univerzitn\'{a} 8215/1, \v{Z}ilina, Slovakia} \address[2]{University Science Park, University of \v{Z}ilina, Univerzitn\'{a} 8215/1, \v{Z}ilina, Slovakia}
[type=editor, auid=000,bioid=1, role=Co-ordinator, orcid=0000-0002-3279-4218] \credit{Conceptualization, Methodology, Investigation, Resources, Supervision} \address[3]{Department of Engineering, Durham University, Lower Mountjoy, South Road, Durham, DH1 3LE, UK} \address[4]{Durham Energy Institute, Durham University, South Road, Durham, DH1 3LE, UK} \address[5]{Institute for Data Science, Durham University, South Road, Durham, DH1 3LE, UK}
[type=editor, auid=000,bioid=1, role=Co-ordinator, orcid=0000-0001-5525-9974] \credit{Conceptualization, Resources} \address[6]{ElaadNL, Utrechtseweg 310 (bld. 42B), 6812 AR Arnhem (GL), The Netherlands}
[type=editor, auid=000,bioid=1, role=Co-ordinator, orcid=0000-0002-5410-6762] \credit{Conceptualization, Methodology, Investigation, Software, Project administration, Resources, Supervision, Visualization, Writing -- original draft, Writing -- review & editing} \address[7]{Department of International Research Projects - ERAdiate+, University of \v{Z}ilina, Univerzitn\'{a} 8215/1, \v{Z}ilina, Slovakia} \fntext[1]{Corresponding author.}
The European Union (EU) is moving towards commitments adopted under the Paris Agreement by aiming at domestic $CO_2$ cuts of at least $40$% below $1990$ levels by $2030$ JRC_web. To tackle this challenge, the deployment of plug-in hybrid electric vehicles (PHEVs), battery electric vehicles (BEVs) and fuel cell electric vehicles (FCEVs) appears to be inevitable JRCreport_2019. Electric mobility is growing at a rapid speed. In $2018$, the number of new electric car sales almost doubled compared to $2017$, and the global electric car fleet exceeded $5.1$ million globalEVOutlook_2019. The world's largest electric car market is the People's Republic of China, followed by the EU and the United States (US). The global leaders in terms of electric car market share are Norway, Sweden, and the Netherlands, which is also dominant in terms of the charging infrastructure density, not only in Europe but also compared to the rest of the world globalEVOutlook_2019. In $2018$, the global EV fleet consumed $58$ TWh of electricity, which is comparable to the total electricity demand of Switzerland in $2017$ globalEVOutlook_2019. The majority of outlooks envisions growing trend. In $2030$, the global electric car sales are expected to reach $23$ million, the stock will exceed $130$ million vehicles (excluding two/three-wheelers), and electricity demand from EVs is estimated to reach almost $640$ TWh globalEVOutlook_2019. In this scenario, slow chargers, which can provide flexibility services to power systems, are estimated to account for more than $60$% of the total electricity consumed globally to charge EVs.
The recent developments in the deployment of electric vehicles are primarily affected by public policies implemented via economic stimuli and technological restrictions addressing industry, consumer behaviour and research activities Bai17. These efforts aim at influencing the course of events in the desired direction. As such, this is a complex process associated with a large number of decision-making situations and potentially leading to unexpected externalities. Hence, it is vital to have at hand tools which are able to address a broad range of decision problems linked to the deployment of EVs on a large scale.
Although electric vehicles have a longer history than fossil fuel vehicles, their mass adoption has started only recently Chan_2013. Initial barriers, such as range anxiety, low availability of charging infrastructure, high price and short lifetime of batteries, are slowly disappearing thanks to technological progress and policies promoting EVs. The impact of policies and the development of methods with the potential to facilitate the EV adoption has been a subject of intense research Bai17,Xie17,Huang18,Sovacool18,Agugliaro17. Specific methodologies have been developed to monitor and assess the behaviour of EV markets and EV components markets Nilsson15,Obrecht19,LaClair19. EV initiatives are largely motivated by sustainability issues. Hence, a reliable assessment of environmental benefits brought by EV adoption is of great importance Hao17,Lave11. Predicted implications of EV charging on the operation of power systems vary across the literature Keoleian12,Sadeghianpourhamami_2018,Muratori18. While some studies highlight positive impacts of reducing the peak power Gonzalez18, others report potentially negative effects, e.g. the level of detected load variability was concluded to be so high that it could potentially limit the integration of EV charging and renewable energy sources Munkhammar18. Underdeveloped charging infrastructure is another major factor hampering large scale EV adoption. For the past few years, the optimal placement and the sizing of EV charging stations have attracted significant attention of researchers Dziurla14,Dou18,Zhao15. The classification of available approaches is given in Sanchari_2018. Charging of EVs is not only influencing and influenced by the integration of renewable energy Gerbaulet15,Qiu16, but it enables bidirectional energy flows Crawford18, and it is expected to affect and be affected by energy markets Samuelsen16 or by the way how EV drivers are navigated Zhou19.
With the growing number of EVs on roads in the last few years, a lot of research in the EV domain has been based on data-centric (or data science) methods and builds on available EV charging data. A recent review Pevec_2019 provides an overview of data sources and summarizes data science literature in the domain of EV charging.
A few data-oriented studies have addressed the charging behaviour of EV drivers. A set of key performance indicators characterizing utilization of chargers was defined and used to compare two rollout strategies: demand-driven and strategic rollout Helmus_2018. No rollout strategy is favourable over the other on all metrics, and the difference between strategies reduces as the EV adoption progresses. Using the location features, i.e. features characterizing the environment in which the infrastructure operates, authors in Straka_2019 built prediction models for the popularity of charging infrastructure (i.e. the number of unique users). The multinomial logistic regression was used by Wolbertus_2018 to determine key factors explaining heterogeneity in the charging duration of categorized charging sessions. The time-of-day-related variables and the type of charging station have the most substantial effect. Some location features such as type of the urban area, the density of chargers and parking possibilities were considered in this study as well. The ability of machine learning methods (random forest, gradient boosting and XGBoost) to predict the idle time (i.e. the time when an EV is connected to a charger without charging) was evaluated by Lucas_2019. The best model is XGBoost, reaching $R^2$ score of $60.3$%. The most influential features are time-of-day-related features and the total energy supplied. Only one location feature, namely the type of the closest road segment, was considered.
Analyses of charging infrastructure utilization focus on temporal aspects, typically applying time-series forecasting. Predictability of energy consumption on single chargers was investigated in Bikcora_2016, finding potentially useful results only for some chargers. Ref. Louie_2017 identified and evaluated the time-series seasonal auto-regressive integrated moving average (ARIMA) models of EV load aggregated over $2400$ chargers. The long-term models (for two years) were found decidedly less accurate than the near-term models (for the most recent 60 weekdays and 24 weekend days). Comparison of ARIMA models with decision trees, considering some exogenous features, concluded that former models are a better choice for forecasting aggregated EV charging loads Buzna_2019. Considering four commonly used machine learning algorithms (K-nearest neighbour, pattern sequence-based forecasting, support vector regression and random forest), forecasts of EV charging load based on customer profile and charger measurements were compared by Gadh16, yielding similar prediction errors. In ai2018household, the day-ahead EV charging load is forecasted as EV charging occurrence-time and the “no charge” day respectively, by several widely used machine learning algorithms. The best performance achieved the hybrid model combining Random Forest, Naive Bayes and XGBoost.
Spatial analyses of charging infrastructure utilization are less developed in the literature than temporal. A preliminary exploratory analysis of spatial patterns formed by energy consumption on charging stations was presented in Lucas_2018. A study found a heterogeneous pattern, observing higher energy intensity in a small number of urban areas and $50$% of the energy supplied comes from $19.6$% of chargers. Ref. Pevec_2018 employs XGBoost model and five features describing the location and parameters of the charging infrastructure to predict its utilization. It also demonstrates how the model can be used to support decisions on locating the charging infrastructure at the level of zones with a radius of $3$ km.
In this paper, we develop a large scale spatial analysis of the energy consumption induced by charging of EVs. Our primary contribution is in extending the available knowledge of location factors, potentially affecting the energy consumption at public charging infrastructure. The candidate set of location factors is large. From the collected data, we extracted more than 120 features. We employ available statistical methods, specifically designed to explore a large set of potentially influential features. We focus on the slow charging infrastructure, which is believed to play a dominant role in the future globalEVOutlook_2019. As a case study, we consider the Netherlands. Recently, several prescriptive models appeared in the literature Pevec_2018,Straka_2019,Wolbertus_2018,Lucas_2019 explaining some performance indicators of charging infrastructure from location features. For example, in Pevec_2018 two such features are used, the number of points of interest and the number of competitive charging stations, to explain the utilization of charging infrastructure. In general, the decision which data to collect when developing a model is difficult. Our secondary contribution is in providing guidelines for these decisions when building prescriptive models involving energy consumption. Our last contribution is methodological. In the analysis, we consider the influence of multicollinearity and statistical stability of selected regression coefficients to improve the statistical reliability of results. In these ways, we contribute to methodologies for decision making support in domains of charging infrastructure deployment and power grids planning.
Charging of electric vehicles is a new field and various terminologies can be found in the literature. We follow Ref. EVdefinitions, where a connector is defined as a physical interface between an electric vehicle and charging infrastructure through which electricity is delivered. Due to the incompatibilities of connector types used by different car manufacturers, several connectors might be available at a charging point, however, no more than one connector can be active at a time. A charging point is an energy delivery device equipped with one or more connectors. The charging point can charge an EV with a power which is less or equal to the maximum power, given in kW, referred to as the charging capacity. A charging station is composed of one or multiple charging points. A user identification interface and all human-machine interfaces are attributed to the charging station and are shared by all charging points. A charging pool is one or a collection of charging stations including the adjacent parking lots. Components of the charging infrastructure are visualised in Figure (ref). A charging transaction starts by plugging a connector into an EV at a start time and ends by unplugging the EV at an \textit{end time}. The difference between end time and start time is \textit{the connection time}. A part of the connection time when EV was charging is referred to as \textit{the charging time}. \textit{The idle time} is the connection time minus the charging time.
The EVnetNL dataset has been provided to us for research purposes by ElaadNL, a Dutch research organisation involved in the development, deployment and operation of EV charging infrastructure. The data is organised in two tables. The table Transactions contains $1~060~763$ rows, each characterising an individual charging event, with columns such as an identifier of a charging point, GPS coordinates of charging station, start time, end time, connection time, idle time, charging time, number of the used RFID cards, consumed energy and unique identifier of the charging event. The second table, Meterreadings, has $32~440~911$ rows, each corresponding to a meter reading that is taken each $15$ minutes if a vehicle is connected. A meter reading is described by the identifier of the transaction, identifier of the charging point, UTC timestamp and value of the meter.
The EVnetNL dataset covers $1747$ charging stations equipped with $2893$ charging points (identified by unique labels) that were operated by the ElaadNL in the period from January 2012 until March 2016. Some charging stations are located close to each other (e.g. in one parking lot). Consequently, the urban context of charging stations cannot be distinguished. Therefore, we considered the charging pools as the main object of the analysis. Locations of charging pools we estimated by aggregating the charging stations. For each charging station, we identified a set of neighbouring stations located within the radius of $50$ meters. For every pair of stations distant more than $50$ meters from each other, we found an empty intersection of sets of neighbouring stations. Therefore, we selected one representative charging station in each set of neighbouring stations, and all stations in the set were merged and formed a charging pool.
In the analysis of energy consumption, we consider only the period from January 1, 2015, until December 31, 2015, as this is the latest available complete year, when the number of charging pools was reasonably stable (see Figure (ref)A). As shown in Figure (ref)B, the number of active users, estimated from the number of RFID cards in use, exceeded the value $15~000$ and it has been relatively stable throughout the year 2015 as well. We obtained a set of $1604$ charging pools operational in $2015$ while being distributed across the entire area of the Netherlands (see Figure (ref)C). In Figure (ref)D we analyse spatial representativeness of the EVnetNL dataset by calculating the ratio between the number of EVnetNL charging pools operational in 2015 and the number of charging pools in the Charging pools 2015 dataset (for more information about the dataset please refer to the Section (ref) of the SI file) for cells of a regular square grid. The ratio takes higher values in the east and south of the Netherlands. The EVnetNL dataset covers smaller cities better, while in large cities, such as Amsterdam and Rotterdam, only a small percentage of charging stations is covered.
We excluded $40$ charging transactions for which either meter values or charging start and stop times were inconsistent across Meterreadings and Transactions tables. To eliminate charging pools with sparse usage patterns, we excluded from the analyses $218$ pools with either less than $30$ charging transactions in 2015 or less than $1$ kW charging capacity. To minimize the effects of the transition period that follows the introduction of a new charging pool, we consider only charging pools that have been in use before January 1, 2015 (we excluded $87$ pools established in 2015 and later). After these rearrangements, we retained $369~550$ transactions taking place on $1~386$ EVnetNL charging pools. The large majority of charging pools possess $1$ or $2$ charging points and deliver power ranging up to $12.5$ kW. Often, the fast charging is declared when the power is exceeding the value of $22$ kW EVdefinitions, hence, all considered charging pools are used for slow charging (see Figure (ref)A-B ).
To characterize the area and human activities taking place in the vicinity of charging pools distributed across the territory of the Netherlands, we collected potentially relevant publicly available geospatial datasets illustrated in Figure (ref).
The geospatial datasets describe locations on the Earth's surface by geometric objects (points, polylines, polygons) and associate geometric objects with alphanumeric attributes. We inspected all available attributes and if multiple datasets contained the same or very similar data, we considered a more complete source or the source providing data with higher resolution. Moreover, we excluded attributes that lead to multicollinearity (e.g. from the triplet of attributes the total population, the population of men and the population of women, we considered only the first attribute as the populations of men and women tend to be very similar, hence, constitute one half of the total population). In what follows, we briefly describe each geospatial dataset. The complete list of selected attributes for each dataset is given in the SI file, Tables (ref)- (ref).
The population cores are continuous spatial units with at least $25$ homes or $50$ registered residents popcores. The dataset associates the population cores with the information about the households (e.g. size and composition), the cardinality of population age groups, and family and civil status of residents. It also includes the information about the employment rate of residents, type of their occupation and characteristics of real-estate properties. The spatial resolution of this dataset is rather low (typically, a population core corresponds to a municipality) and we use it to investigate whether aggregate characteristics of municipalities can explain the usage pattern of charging pools. From this dataset, we selected 45 attributes for further analysis (see Table (ref) of the SI file).
The neighbourhoods are spatial units that are approximately uniform when considering the type of built-up area or the socio-economic indicators neighbpop. The neighbourhoods are used by the Dutch national statistical office to collect, maintain and distribute the statistical data. In addition to the population data, the Neighbourhoods dataset maintains information about the number of address points within the neighbourhoods, level of the income of the residents, number of registered private and company vehicles and so on. From the available attributes, we selected 63 attributes listed in Table (ref) of the SI file.
The Energy consumption dataset contains records of the annual natural gas and electricity consumption in residential houses and industrial facilities, together with the number of buildings equipped with metering devices energyatlas. The spatial resolution is identical with the Neighbourhoods dataset. For further analyses, we selected 12 attributes that are detailed in Table (ref) of the SI file.
The Liveability dataset has been introduced by the Dutch Ministry of the Interior and Kingdom Relations to monitor the quality of living in Dutch neighbourhoods liveability. We use the liveability index 2016, that was revised in 2015. The liveability is quantified by a composite index and by five specific indices evaluating categories such as housing, socio-economic background of residents, services, safety, and the environment. Hence, from the liveability dataset, we extracted 6 attributes listed in Table (ref) of the SI file.
We use the LandScan 2015 landscan high resolution population raster grid estimating the 24-hour average of population count with a spatial resolution of approximately $1~$km $\times~1~$km. In contrast to the Population cores or the Neighbourhoods datasets that capture the residential population only, the Landscan considers the mobility of residents.
The Land use dataset describes the occupation of land in the Netherlands by polygons lccbs. Each polygon is assigned an attribute value taking one out of predefined classes of land use. Examples of land use categories are traffic areas,building sites, recreational areas and business areas. Complete list of 25 categories is given in Table (ref) of the SI file.
To model the impact of traffic on the usage of charging pools, we consider the traffic flows dataset traffic_flows. The dataset is organized around a high resolution model of the road network and the traffic flow information is added in the form of attributes that are associated with the road segments. The description of 9 attributes that have been selected to compile features is given in Table (ref) of the SI file.
The OpenStreetMap (OSM) is one of the most successful free maps osm. From the OpenStreetMap of the Netherlands, we extracted all points of interest (POIs) considering $2$ km $\times~2$ km squared areas centred at the positions of the EVnetNL charging pools. We identified $593$ different POI types, some of them appearing in only very few instances. For this reason, we associated manually POI types with one of the 15 categories listed in Table (ref) of the SI file. The POIs, organized in these $15$ categories, were used to model the venues in the proximity of charging pools that are often visited by EV drivers.
Aiming at estimating the positions of all available charging pools present in the Netherlands by the end of the year 2015, the Charging pools 2015 dataset was compiled from the EVnetNL, OpenChargeMap ocm and OplaadPalen opp datasets. Utilizing the date when a charging station was added to the dataset, we extracted from the OpenChargeMap and OplaadPalen positions of all charging stations that were available by the end of 2015. As with the EVnetNL dataset, we estimated the position of charging pools from the positions of charging stations, while utilizing the information about the geographical proximity. In the first step, we added to the Charging pools 2015 all EVnetNL charging pools available in 2015. In the second step, we added one-by-one to the Charging pools 2015 dataset the charging stations from the OpenChargeMap and OplaadPalen datasets, if their position was more than $50$ meters distant from already added pools. This way, we obtained positions of $8~366$ charging pools (see Figure (ref)C).
To prepare the data for the analyses, we applied the pre-processing procedure composed of three stages: the missing value handling, the extraction of features and the analysis of potential data modelling problems, described in the following subsections.
Values of some attributes associated with geometric objects in geospatial datasets are missing. By visualising the geometric objects with missing attribute values on the map, we identified two main sources of problems. Some geometric objects with missing data represent water landscape. In such a case, we excluded the geometric objects from datasets. The second source of specific problems with missing data are cities of Baarle-Nassau and Baarle-Hertog with peculiar borders between the Netherlands and Belgium. In this case, we ignored missing values as the EVnetNL dataset does not feature any charging pool in this area. We applied a two-step approach to handle missing values (see section (ref) of the SI file). In the first step, when it was possible we estimated the missing values from the available data. Further analysis revealed that in the geospatial data the attribute values tend to be missing in areas with low intensity of human activities (e.g. low population, low density of buildings, low electricity consumption, etc.). Therefore, in the second step, we applied simple rules that set the missing values of some attributes to zero (or lowest possible value) in areas with low intensity of human activities. Remaining missing values are addressed after generating features characterising the vicinity of charging pools.
We modelled each charging pool as a single point defined by GPS coordinates. We used a buffer, circular area centred at the position of charging pool and having the radius $r$, to model the vicinity of charging pools. Values of features were calculated from GIS polygon data, considering spatial intersections between the area of buffers and GIS geometric objects, while assuming a uniform spatial distribution of considered quantities over the area represented by GIS geometric objects. For POI data (OpenStreetMap and Charging Stations 2015), the features are defined two ways: as the distance from the pool to the closest point of interest and as the density of points of interest within the buffer area. We set the buffer radius to $r~=~350$ meters. For more details please refer to Section (ref) of the SI file.
To ensure that the handling of missing values cannot influence significantly the results of the analysis, we applied to each feature the following rule: The feature is used in the analysis if areas with missing values of the attribute take less than $15$% of the buffer areas, otherwise it is excluded. To make sure that features are not built based on GIS data with a large proportion of missing values, if there was less than $1.5$% of feature values missing after applying the estimations, missing values were imputed by a median value, otherwise, the feature was discarded. Finally, we obtained $\boldmath{195}$ features.
Adopting the notation from Ref. hastie2015statistical, features are organised in a matrix $\bm{X} = (\bm{x}_1,~\dots~,\bm{x}_p)$, where the column $\bm{x}_j$ corresponds to the feature $j~=~1,~\dots,~p$ and $x_{ij}$ is the $i$-th observation of the feature $j$, for $i~=~1,~\dots,~n$, where $n$ is the number of observations (charging pools). The response vector $\bm{y}$ represents the energy consumption in kWh on charging pools during the year 2015, i.e. $y_i$ is the energy consumption of the charging pool $i$.
When a dataset is fit with a regression model, many problems may occur. Several steps, recommended in the literature James_2014,Kuhn_2013,Afifi_2011, were applied to the feature matrix $\bm{X}$ and to the response vector $\bm{y}$, to explore and to address potential problems. The sequence of steps is illustrated in Figure (ref).
First, uninformative features, i.e. those with more than $95$% of zero values, were excluded. Dependencies between features can cause serious problems when interpreting results obtained by regression methods, referred to as collinearity and multicollinearity. Despite the potential to bias the interpretation of results, these problems tend to be overlooked in the data analysis hsu2015identifying. We identified groups of features with absolute value of the Pearson correlation coefficient between each pair of features in the group greater than $0.95$ Kuhn_2013. For each group, we chose one feature as a representative of the group. The selected feature was included, while other group members were excluded from further analysis. Identified groups and selected representative features are listed in Table (ref) of the SI file. To mitigate multicollinearity, we applied a procedure composed of two steps: in the first step, the values of the variance inflation factor (VIF) James_2014 were calculated. In the second step, we identified feature with the maximum value of the VIF. If the maximum VIF was greater than or equal to value $10$, the corresponding feature was excluded from further analysis and the procedure was repeated, otherwise, the procedure was terminated. The list of excluded features is provided in Section (ref) of the SI file. These steps reduced the multicollinearity to the level recommended in the literature, while quantified not only by the values of the VIF but also by values of measures derived from eigenvalues of the correlation matrix chatterjee2015regression.
To explore whether a nonlinear function of $\bm{X}$ could explain response vector $\bm{y}$ better than the linear function, we transformed original features $j~=~1, \dots, p$, except binary features, using functions $\sqrt{\bm{x}_j}$, $\bm{x}_j^2$ and $\log(\bm{x}_j+1)$ and the feature matrix $\bm{X}$ was extended by transformed features (in this section, functions are applied to vectors element-wise). By combining the ordinary least squares (OLS) with 10-fold cross-validation James_2014, the extended feature matrix was fit to the response vector $\bm{y}$ and the mean squared error (MSE) was evaluated on the test data. From the results, we concluded that the basic non-linear transformations do not improve the fit sufficiently to compensate for additional model complexity and for making the model more likely to overfit the data. Hence, we do not apply any non-linear transformation to the feature matrix $\bm{X}$. Furthermore, we evaluated several transformations of the response vector: $\sqrt{\bm{y}}$, $\bm{y}^2$, $\log(\bm{y})$ and the Box-Cox transformation Kuhn_2013. By analysing the residual plots (see Figure (ref) of the SI file) we concluded that the transformation $\log(\bm{y})$ significantly improves the fit and we apply it in the regression analyses when fitting the energy consumption with the feature matrix, presented in Section (ref).
Finally, we obtained $1259$ observations and $119$ features derived from geospatial datasets characterizing the vicinity of charging pools and $5$ features derived from EVnetNL dataset specifying the location and basic characteristics of charging pools.
This paper aims at explaining consumed energy on charging pools using features derived from the geospatial datasets. This task can be formulated as a regression problem combined with the statistical inference, indicating the statistical strength of features. From the geospatial datasets, we extracted a relatively large number of features. To facilitate the interpretation of results, features that have a higher potential to explain response variable shall be selected. We assume that from all features, only a relatively small number plays an important role. Hence, we use the Lasso method Tibshirani96, which combines the parameter fitting with the variable selection functionality.
Considering the collection of $n$ samples $\{(\bm{x}_i, y_i) \}_{i=1}^{n}$, for some $\lambda \geq 0$ the Lasso method solves optimization problem
where the scalar $\beta_0$ (intercept) and vector $\beta$ (regression coefficients) are optimization variables. The first term corresponds to the least squares objective function and its role is to ensure a good fit between the linear regression model $\eta(\bm{x}_i) = \beta_0 + \sum_{j=1}^{p} x_{ij} \beta_j$ and response value $y_i$, while the second term regularizes the estimated values of regression coefficients in a way that leads to a variable selection (i.e. some regression coefficients are set to zero). The parameter $\lambda$ determines the trade-off between goodness of the fit and strength of the variable selection. Often, the factor $\frac{1}{2n}$ in ((ref)) is replaced, with $1/2$ or $1$. Although this corresponds to a simple reparametrization of $\lambda$, the factor $\frac{1}{2n}$ makes $\lambda$ values comparable across different sample sizes, which is useful for cross-validation hastie2015statistical. The solution of problem ((ref)), $\hat{\beta_0}$ and $\hat{\beta}$, constitutes the estimate of the model parameters.
The Lasso method tends to have some difficulties with the identification of relevant features on datasets with highly correlated features hastie2015statistical. The Elastic Net method, i.e. adding the term $\lambda_2 \sum_{j=1}^{p} |\beta_j|^2$ (with $\lambda_2 \geq 0$) to ((ref)), may help to identify correlated features hastie2015statistical. In numerical experiments, the Elastic Net method was tested as well, however, it selected largely the same set of features as the Lasso method.
The parameter $\lambda$ in ((ref)) controls the complexity of the model. A smaller value of $\lambda$ results in a larger number of non-zero regression coefficients and allows the model to adapt more closely to data, however, it can lead to overfitting. On the contrary, a larger value of $\lambda$ leads to a sparser and more interpretable model with the risk of preventing the Lasso from capturing the main signal in the data. Hence, the value of $\lambda$ should be carefully chosen. To estimate a suitable value of $\lambda$, we evaluated a range of values using log spacing. For each value, we used the $k$-fold cross-validation hastie2015statistical to evaluate the MSE of the model. Finally, we found the $\lambda^{CV}$ for which the minimum MSE was achieved and we selected the corresponding $\hat{\beta}_0^{CV}$ and $\hat{\beta}^{CV}$.
Traditional methods, such as OLS, determine the statistical strength of features by evaluating $p$-values. The results obtained by the OLS regression on features selected by the Lasso method cannot be fully used in post-selection analysis as the exclusion of some features causes a bias hastie2015statistical. The adaptive nature of the Lasso method makes the problem of estimating $p$-values difficult--both conceptually and analytically. From the available approaches to the inference problem, we selected the bootstrap hastie2015statistical. The bootstrap was recommended as a suitable method for the assessment of the stability of selected regression coefficients in hsu2015identifying, even though there it has not been used in the analysis while taking a risk of presenting an unreliable set of significant coefficients. The bootstrap is a generic tool for assessing the statistical properties of complex estimators. First, the dataset is sampled. Second, the $k$-fold cross-validation is applied to each sample, to find $\lambda^{CV}$, $\hat{\beta}_0^{CV}$ and $\hat{\beta}^{CV}$. The distributions of $\hat{\beta}^{CV}$ coefficients obtained across all samples and the counts of samples resulting in coefficients $\hat{\beta}^{CV}$ equal to zero are used to assess the significance of features hastie2015statistical
We prepared and modelled the data using the R language. We processed the GIS data with sf, raster and osmar packages. The distribution of energy consumption was fitted with the fitdistrplus package. The Lasso method, including model selection, is implemented in the glmnet package. For the $k$-fold cross-validation we used $k$ = 10. We considered the values $10^i$, for $i$ in the interval from $-4$ to $0$ in steps of $0.02$, when applying the cross-validation to explore the values of the parameter $\lambda$ in the Lasso method. When studying the stability of coefficients selected by the Lasso method, we used the bootstrap method with $10~000$ realisations.
Charging pools differ in the maximum capacity and in the number of charging points (see panels A and B of Figure (ref)), potentially leading to differences in the consumed energy. As illustrated in Figure (ref)A, charging pools with a higher number of charging points tend to have slightly higher energy consumption. Moreover, the number of cases when two transactions run on a charging pool in parallel is not negligible (see inset of Figure (ref)A). To analyze to what extent it is important to account for these effects, we define several simple metrics of the consumed energy at charging pools in Table (ref).
To gain some basic understanding, whether some measures can be better explained with the feature matrix, we ran the OLS regression. In Table (ref) we report the values of the coefficient of determination, $R^2$. The highest $R^2$ we obtained for the consumed energy at a charging pool. High similarity in $R^2$ values we attribute to the high correlations between response vectors. The Pearson correlation coefficient for all pairs of response vectors ranges from $0.86$ to $0.98$. Hence, for further analyses, we chose the consumed energy at a charging pool as the response vector $\bm{y}$.
The density histogram indicates that the consumed energy at charging pools is rather heterogeneous and positively skewed (see Figure (ref)B). We observe high consumption on a small number of charging pools and small consumption is observed for a large group of charging pools. To model the distribution of energy consumed at charging pools, we considered original data and five transformations ($\bm{y}^2$, $\bm{y}^3$, $\sqrt{\bm{y}}$, $\sqrt[3]{\bm{y}}$ and $log(\bm{y})$) and parametrized density functions of three known probabilistic distributions (Weibull, beta and gamma). We used the Kolmogorov-Smirnov goodness of fit test together with the inspection of the P-P and the Q-Q plots warren2010application to conclude that the transformation $\sqrt[3]{\bm{y}}$ combined with the beta distribution provides the best fit to the data. The fitting procedure is detailed in Section (ref) of the SI file. The functional form of the probability density function is derived in Appendix (ref) and takes the form
The symbol $B(\alpha, \beta)$ denotes the beta function and the estimates of parameters take the following values: $\alpha = 2.58$ , $\beta = 4.53$, $y_{min} = 91.55$ kWh and $y_{max} = 16~649.40$ kWh.
We gained interesting insights by analyzing the relationship between energy consumption and other charging pool indicators constituting the energy consumption. The energy consumption at a charging pool $i$ can be decomposed into the product of three other indicators, i.e.
where $n_i$ is the number of charging transactions taking place at a charging pool $i$, $t_i$ is the average charging time per transaction at a charging pool $i$ and $p_i$ is the average charging power at a charging pool $i$. We organized these quantities for all charging pools as vectors $\bm{n}$, $\bm{t}$ and $\bm{p}$. To assess the role of these three factors, in the heterogeneity of the consumed energy across charging pools, we explored six models presented in Table (ref). Models are based on Eq. ((ref)), where one indicator or product of two indicators, represented by the regression coefficient $k$, is considered invariant across charging pools.
We obtained the estimate $\hat{k}$ of the coefficient $k$ from the data by using the OLS method. Among the models that explain the consumed energy from one indicator, the highest value of $R^2$ is obtained for the model explaining the consumed energy from the number of charging transactions. Similarly, models explaining energy consumption from pairs of indicators that include the number of transactions have high values of $R^2$. Hence, the major factor associated with the heterogeneity in the consumed energy across charging pools is the number of transactions. The fluctuations in the charging patterns (i.e. average charging time and average charging power) play a much smaller role.
In Table (ref), we calculated, mean, standard deviation and the coefficient of variation over all charging pools for quantities which are represented by the coefficient $k$. For example, in the model $\bm{y} = k\bm{n}$, the coefficient $k$ that replaces in Eq. ((ref)) the expression $t_i p_i$ can be interpreted as the average of consumed energy per transaction. Hence, the average consumed energy per transaction is $8.69$ kWh, the charging time per transaction takes on average $2.52$ hours and the average power reaches $3.45$ kW. These numbers indicate that the majority of charged EVs are plug-in hybrids. The charging capacity, which is for the majority of charging points around $11$ kW (see Figure (ref)), is significantly underused. The values of the coefficient of variation are high when the number of transactions is included in the analysed quantity, confirming that the largest variance is associated with the number of transactions.
As the number of charging transactions is closely associated with the consumed energy at charging pools, it is crucial to understand the way EV drivers decide which charging opportunity they choose. We hypothesize that some exogenous factors, characterizing the environment surrounding the charging pools affect the consumed energy at charging pools and we explore it in the next section.
We applied the methodology described in Section (ref) to the feature matrix derived from geospatial datasets, first without considering the features derived from the EVnetNL dataset (see Section (ref)), and to the $log$-transformation of the response vector $\bm{y}$, representing the consumed energy at charging pools. To facilitate the mutual comparison of regression coefficients, we standardized each coefficient by multiplying it with the standard deviation of the bootstrap sample of the predictor. Thus, the standardized coefficient can be considered as an estimate of the change in the response variable, when the feature increases by one standard deviation. The larger is the absolute value of the median of standardized regression coefficient samples, the stronger is the feature's potential to influence the response variable [p. 372]siegel2016practical. Hence, the absolute value of the median of bootstrap realisations can be considered as a measure of the feature strengths. The different signs of regression coefficients, across bootstrap realisations, can be attributed to the low significance or to the simultaneous selection of correlated features hastie2015statistical. Hence, the more consistent are the values of standardized regression coefficient across bootstrap samples, the more significant is the feature corresponding to a regression coefficient. As a rule of thumb, we consider as significant those features where the number of bootstrap samples with zero coefficient value is less than $5\%$ and the number of samples with the opposite coefficient sign to the sign of the majority of the sample is negligible. To provide a broader view on results, we show in Figure (ref) the empirical distributions of standardized regression coefficients that reached the value of zero in maximum $10\%$ of the bootstrap realisations. The statistical inference is sometimes omitted in energy studies and only a single run of a variable selection method is evaluated hsu2015identifying, ma2016estimation. The inspection of Tukey's boxplots justifies the use of the bootstrap or in more general, the use of a statistical inference method. The majority of presented coefficients obtained a zero value in few samples or exceptionally the opposite sign, that could lead to omitting these coefficients or interpreting them in a wrong way if evaluating only a single run of the Lasso method.
We show the regression coefficients in descending order of median value in Figure (ref). To clarify the results, we organized significant features with the positive sign of the median in four groups:
Similarly, we organized significant features with the negative value of the median into a group:
The largest number of significant features we found in the population group, which is partly because this group contains most of the features. Many significant features point to a single factor. The largest group of features, $PC_{32}$, $PC_{29}$, $N_{36}$, ($N_{55}$), indicate that high (low) income and wealth (poverty) are positively (negatively) linked with the amount of consumed energy at charging pools. Most likely, this is due to the high prices of EVs, making them more affordable for better-situated residents and businesses. Probably for the same reasons, some significant features ($PC_{11}$, $N_{13}$) are linked with children or youth and to the elderly or retired population ($N_{15}$), i.e. social groups that are typically less wealthy. The high percentage of vacant houses ($N_{38}$) and the high average value of residential real estate property ($N_{36}$) are associated with high energy consumption at charging pools. It could correspond to newly built and yet not entirely inhabited areas, with a higher standard of living expressed in higher real estate values. Note, if a feature representing a minimum distance to an object has a positive coefficient, then the energy consumption increases with increasing the distance from a given object and vice versa. Hence, the proximity of hobby related points of interest is negatively associated with the energy consumption at charging pools. Furthermore, the density of roads is also among features that are positively linked with the energy consumption at charging pools, indicating that good access to charging pools contributes to energy consumption.
We included to the feature matrix some features (number of charging points, maximum power, latitude and longitude and the rollout strategy) derived from the EVnetNL dataset (see Section (ref)) and repeated the analysis. The results are shown in Figure (ref) of the SI file. In general, significant features are similar as in Figure (ref); however, the number of significant features is slightly smaller. The significance of some features from groups Physical environment(+) and Population(-) was reduced. All EVnetNL features are significant, while the maximum charging capacity, the \textit{number of charging points} and \textit{longitude} have high strength, indicating that parameters of charging pools potentially influence the energy consumption. We attribute the reduction in the number of significant features to the replacement of some features by EVnetNL features. For example, previously significant feature, \textit{open wet natural terrain and water} ($LC_{24}$), seems to be expressed via the \textit{longitude}. The negative influence of \textit{longitude} can be explained by the geography of the Netherlands, whereas the western part of the country is more urbanized and we find here the majority of large Dutch cities. At the same time, there is a lot of surface water as the western part of the country is largely situated below the sea level.
The majority of charging pools has been located using one out of two (strategic or demand-driven) rollout strategies Helmus_2018. The strategically located charging pools are placed near public facilities, where the EV charging is intuitively expected. The demand-driven charging pools are built upon the request from EV users, typically near to their homes. In this section, we investigate whether the rollout strategy makes a difference in factors associated with energy consumption. Information about the rollout strategy is available in the EVnetNL dataset and we used it to split charging pools into two groups. We applied the Lasso method separately to each group considering the feature matrix without the features derived from the EVnetNL dataset. The selected features in Figure (ref) coincide to a large extent with factors selected for the complete dataset (see Figure (ref)), however, now we can observe differentiation of factors according to the location strategy.
The energy consumed at strategically located charging pools, (Figure (ref)A), is positively linked to the working sector of residents and the physical environment, i.e. to certain types of venues adjacent to the charging pool. Working sectors of residents represented by $PC_{32}$, $PC_{33}$ and $PC_{34}$ indicate the prevalent businesses in municipalities, positively associated with energy consumption. Moreover, selected features for strategical rollout refer to some businesses and some locations (sports fields, socio-cultural facilities) that could be associated with occasional charging.
For the charging pools with the demand-driven rollout, (Figure (ref)B), the negative coefficients of deaths per 1000 residents and live births per 1000 residents indicate that areas with higher natality and mortality are negatively linked with the energy consumption, pointing out to areas inhabited by socially weaker groups of specific age categories.
We have tested several other stratifications, e.g. based on the number of charging points, the proportion of the residential area within the charging pool's buffer, the administrative division of the Netherlands into provinces and the number of residents of municipalities. Except for the last criterion, we obtained only a very small number of selected features. Dividing charging pools into two groups based on the municipality population, considering a threshold of $50~000$ residents, we obtained approximately the same size of groups. Interestingly, we observe that charging pools located in municipalities with more than $50~000$ residents consume more energy on average by $48$% than charging pools in the other group of municipalities. We find the higher number of significant features, linked to municipality population characteristics, financial and real estate businesses and the physical environment, for the charging pools located in municipalities with a smaller population (see Figure (ref) in the SI file).
We analyzed the explainability of consumed energy at charging pools from several points of view. Main conclusions derived from the data analysis are the following:
Our results extend the knowledge base about the energy consumption at charging pools and provide an evidence for a relationship between energy consumption and the characteristics of the vicinity of charging pools opening several possibilities for future applications.
The methodology and our findings can be used to fine-tune rollout strategies for deployment of charging pools. A rollout strategy optimized for a specific group of charging pools, e.g. charging pools used for work-charging, can be enhanced by applying the presented methodology to this specific group of charging pools and identified characteristics can be used to select locations of new charging pools in a way which corresponds better to energy constraints. Information on selected regression variables can be used to build prediction models in a more targeted way. The presented results are applicable to the Netherlands, however, the proposed methodology can be used also elsewhere.
The selected regression coefficients can be associated with three different spatial scales. Some describe close vicinity of charging pools, e.g. the number of financial and real estate businesses ($N_{31}$), others are attributed to the municipality, e.g. features derived from the Population cores datasets such as the percentage of employed residents in the municipality working in a non-commercial sector ($PC_{33}$). The last group of regression coefficients has the potential to characterize the location of a charging pool at the country level, e.g. latitude and longitude (see Figure (ref) in the SI file). Hence, while properly considering the spatial level of selected regression coefficients, a hierarchical or a customized rollout strategy, considering specific geographic scales, could be designed.
Apart from enhancing the strategies for charging pools deployment, our study can help to improve energy demand models for power grid capacity planning by better considering the energy demand at charging pools and the proposed methodology could be also applied to other service systems, e.g. to stations for shared electric cars, scooters or bicycles.
Several important limitations are inherited from the used statistical methods. Apart from notorious limitation “correlation does not imply causation”, which requires careful consideration of all results, we wish to point out limitation to the interpretability of our results due to the presence of multicollinearity in the input data. To obtain statistically stable results, we reduced the level of multicollinearity to the level recommended in the literature. However, the removal of some features hampers the interpretation of our results. Similarly, it is likely that some relevant data representing important determinants of energy consumption are missing in the analyses, e.g. mobility behaviour of the population or visitation patterns of venues located in the vicinity of charging pools, as we were not able to collect them. Further research could focus on a more complex characterization of the usage of charging pools, specific groups of charging pools e.g. determined based on similarities in usage patterns, or exploration of possibilities to extend the proposed methodology and findings to other geographic areas.