EconBase
← Back to paper

A Data Fusion Approach for Ride-sourcing Demand Estimation: A Discrete Choice Model with Sampling and Endogeneity Corrections

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,081 characters · 19 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\thispagestyle{empty}

spacing{1.2} \begin{center} A Data Fusion Approach for Ride-sourcing Demand Estimation: A Discrete Choice Model with Sampling and Endogeneity Corrections \\ \end{center} \begin{flushleft} 14 October 2022 \\ Rico Krueger (corresponding author) \\ Department of Technology, Management and Economics \\ Technical University of Denmark (DTU), Denmark \\ [email removed] \\ Michel Bierlaire \\ Transport and Mobility Laboratory \\ Ecole Polytechnique F\'{e}d\'{e}rale de Lausanne, Switzerland \\ [email removed] \\ Prateek Bansal \\ Department of Civil and Environmental Engineering\\ National University of Singapore, Singapore \\ [email removed] \\ \end{flushleft}

\thispagestyle{empty}

Abstract

Ride-sourcing services offered by companies like Uber and Didi have grown rapidly in the last decade. Understanding the demand for these services is essential for planning and managing modern transportation systems. Existing studies develop statistical models for ride-sourcing demand estimation at an aggregate level due to limited data availability. These models lack foundations in microeconomic theory, ignore competition of ride-sourcing with other travel modes, and cannot be seamlessly integrated into existing individual-level (disaggregate) activity-based models to evaluate system-level impacts of ride-sourcing services. In this paper, we present and apply an approach for estimating ride-sourcing demand at a disaggregate level using discrete choice models and multiple data sources. We first construct a sample of trip-based mode choices in Chicago, USA by enriching household travel survey with publicly available ride-sourcing and taxi trip records. We then formulate a multivariate extreme value-based discrete choice with sampling and endogeneity corrections to account for the construction of the estimation sample from multiple data sources and endogeneity biases arising from supply-side constraints and surge pricing mechanisms in ride-sourcing systems. Our analysis of the constructed dataset reveals insights into the influence of various socio-economic, land use and built environment features on ride-sourcing demand. We also derive elasticities of ride-sourcing demand relative to travel cost and time. Finally, we illustrate how the developed model can be employed to quantify the welfare implications of ride-sourcing policies and regulations such as terminating certain types of services and introducing ride-sourcing taxes. \\ \\ Keywords: ride-sourcing, big data, mode choice, endogeneity, travel demand.

Introduction

Ride-sourcing services like Uber, Lyft, Didi and Grab have expanded rapidly in the last decade and have attracted considerable ridership in many metropolitan areas worldwide goletz2021ride. Ride-sourcing is a disruptive transport mode with positive (provision of convenient, affordable on-demand transportation options) and negative (congestion, pollution, increased vehicle kilometres travelled, possible cannibalisation of public transport demand) impacts on transport systems tirachini2020ride. To realise their advantages and inhibit their disadvantages, ride-sourcing services need to be planned, regulated and managed goletz2021ride, tirachini2020ride. To that end, a rigorous understanding of ride-sourcing demand is essential. Specifically, it is crucial to i) explain the characteristics of ride-sourcing demand, ii) analyse the interaction of ride-sourcing with other transport modes and iii) quantify the welfare implications of introducing ride-sourcing services or amending operational policies. To provide actionable, evidence-based decision support, ride-sourcing demand analysis calls for i) powerful methods to leverage datasets with varying disaggregation and resolution and ii) comprehensive datasets with user-level preferences and vehicle-level operations at an urban scale, yet both are currently found wanting.

In terms of methods, ride-sourcing demand analysis is currently dominated by approaches without adequate foundations in microeconomic theory. Aggregate models such as regression models for count and continuous data ghaffar2020modeling, marquet2020spatial estimate a statistical relationship between realised aggregate demand and aggregate explanatory variables. Ordered outcome models for explaining ride-sourcing use at the individual level alemi2019drives, von2021exploring infer structural relationships between demand and individual-specific attributes. These methods disregard that ride-sourcing demand arises at the disaggregate level in the form of a mode choice involving trade-offs between various alternative-specific attributes (e.g. travel cost, time, reliability and safety).

Most applications of these statistical methods are driven by limited data availability. Commonly considered data sources exhibit significant weaknesses when they are analysed in isolation. Household travel surveys have a broad geographical coverage. However, they typically only include a small number of ride-sourcing trips, which precludes a thorough analysis of ride-sourcing demand. In principle, discrete choice experiments (DCEs) allow for a detailed analysis of ride-sourcing demand. However, data collected via DCEs may exhibit hypothetical biases. Also, DCEs are typically not repeated over time due to financial and logistical constraints. Recently, ride-sourcing trip records have been published under data sharing agreements between ride-sourcing companies and city authorities ghaffar2020modeling. These trip records have a broad spatiotemporal coverage. However, in isolation, they cannot be used for the disaggregate analysis of ride-sourcing demand since they do not contain any information about the demand for other modes.

This research aims at improving ride-sourcing demand analysis. To that end, we present and apply an approach for estimating ride-sourcing demand using discrete choice models (DCMs) by fusing multiple data sources. DCMs are well suited for analysing ride-sourcing demand due to their solid foundations in microeconomic theory. DCMs estimate structural relationships between observed travel choices and various alternative- and individual-specific attributes. Due to their structural nature, DCMs produce stable and transferable predictions, which in turn makes DCMs suitable for analysing counterfactual pricing and service configuration scenarios.

We first construct an estimation sample of trip-based mode choices in Chicago by enriching household travel survey data with publicly available ride-sourcing and taxi trip records. By fusing the two data sources, we address the problems of i) having too few ride-sourcing trip records in household travel survey data and ii) having no information about other modes in ride-sourcing trip records. However, constructing an estimation sample from two revealed preference data sources creates two challenges in developing a DCM. First, a sampling correction is needed to account for the enrichment of the household travel survey data with ride-sourcing and taxi trip records. Second, the constructed mode choice dataset likely exhibits endogeneity biases, as the demand for the ride-sourcing options and their prices is simultaneously determined by supply-side constraints and surge pricing mechanisms castillo2017surge. We address both challenges by formulating a multivariate extreme value (MEV)-based DCM with sampling and endogeneity corrections. To correct for sampling biases, we adopt a conditional maximum likelihood estimator bierlaire2020sampling due to its efficiency properties; and to correct for endogeneity biases, we adopt the control function approach petrin2010control due to its simplicity. Ultimately, we apply the model to the constructed mode choice dataset to analyse the demand for ride-sourcing services in Chicago. The parameter estimates of the DCM are translated into the elasticity of ridesourcing demand relative to alternative-specific attributes like price and travel time. We also illustrate how the estimated model can be employed to quantify the welfare implications of ride-sourcing policies and regulations such as eliminating certain types of services and introducing taxes.

We organise the remainder of this paper as follows: In Section (ref), we review related literature. In Section (ref), we describe the construction of the estimation sample for the empirical application. In Section (ref), we present the general formulation of the econometric model. In Section (ref), we explain the model specification considered in the empirical application. In Section (ref), we discuss the results of the empirical application, and in Section (ref), we present the welfare analysis. Finally, we conclude in Section (ref).

Related literature

The literature on ride-sourcing demand analysis evolves rapidly. In Table (ref) in Appendix (ref), we present an overview of recently published ride-sourcing demand analysis studies. For reviews of earlier studies, the reader is directed to tirachini2020ride and wang2019ridesourcing. The studies enumerated in Table (ref) can be subsumed under five topics:

enumerate• travel modes (mostly public transit) that are complemented or substituted by ride-sourcing services; impact of emerging on-demand mobility services on vehicle ownership; • association of built-environment, socio-demographics, weather and land use characteristics with ride-sourcing demand at a spatial level (e.g. census tract and census block groups); • association of attitudes, socio-demographic and economic characteristics with the usage of and preferences for ride-sourcing services; • determinants of preferences for the use of pooled ride-sourcing services; • impact of mode-specific attributes (e.g. travel time and wait time) on the demand for these services in the multi-modal transport system.

The studies use mainly three types of data (see column “data type” in Table (ref)). First, trip records from ride-sourcing companies are merged with supplementary data about land use, weather and census tract attributes. These studies focus predominantly on the first two of the five topics enumerated above; only a few studies focus on the fourth topic. Second, household travel surveys with information about individuals' travel patterns, socio-demographic characteristics and attitudes are considered for exploring the third topic; a handful of studies also focus on topics one and four. Third, DCEs are employed to investigate topics three to five.

In terms of methods, most studies that consider the first data type aggregate trips across space and time and then rely on geographically weighted and spatially lagged count or continuous data models with autoregressive structure or panel effects. A few studies use off-the-shelf machine learning algorithms such as random forest or gradient boosting decision trees. Studies considering household travel survey data use multinomial and ordered logit models. Several studies also develop joint models of continuous, count and ordered variables (i.e.\ generalised heterogeneous data models). Structural equation models are also used to analyse the relationship between ride-sourcing demand and latent attitudes. Finally, studies collecting data through DCEs naturally employ DCMs such as nested, latent class, error component or mixed logit models.

Only a few studies use DCMs and revealed preference data to analyse ride-sourcing demand. habib2019mode considers revealed preference data from a household travel survey to investigate ride-sourcing demand using a semi-compensatory choice model with probabilistic choice set formation. The study finds that ride-sourcing demand mostly complements the demand for driving and transit and substitutes taxi demand. Furthermore, the probability of considering ride-sourcing varies by age, whereby young people are more likely to consider ride-sourcing, and older people are more likely to consider taxis. lam2021geography considered revealed preference data constructed from ride-sourcing trip records, field data and API queries to analyse ride-sourcing demand in New York City. The authors employ an aggregate logit model for market share data. The model includes endogeneity corrections for price and wait times. The study finds that the distribution of ride-sourcing benefits varies substantially across space, with low accessibility areas experiencing comparatively higher benefits.

In summary, household travel surveys and trip records have been used in isolation. Both data sources exhibit significant weaknesses when used in isolation: Household travel surveys contain insufficient information about ride-sourcing demand; trip records cannot be used for disaggregate demand analysis, as they do not include information about individual-level preferences for other travel modes. This current study contributes to the literature with a DCM for disaggregate demand analysis of ride-sourcing services by fusing both datasets and addressing potential endogeneity issues. This data fusion framework leverages the richness of both data types while addressing the shortcomings of analysing them in isolation. We also control for demographics, transit accessibility, parking cost, land use characteristics, pedestrian friendliness, and weather conditions in the DCM, which may affect demand for ridesourcing services. The developed model can be used as a direct input to activity-based travel demand forecasting models to quantify the short- and long-term impact of policies and regulations related to ridesourcing services on the multi-modal transport system.

Finally, our study is also related to the literature on endogeneity and discrete choice analysis in various other applications, including but not limited to consumer choice petrin2010control, residential location choice guevara2012change, airline itinerary choice lurkin2017accounting and parking choice gopalakrishnan2020combining.

Data

We construct an estimation sample of trip-based mode choices in Chicago from November 2018 to May 2019, following the process visualised in Figure (ref). In what follows, we describe the construction of the estimation sample in detail.

figure[figure omitted — 146 chars of source]

Primary data sources

Trip records for the construction of the estimation sample are gathered from two primary sources, namely i) ride-sourcing and taxi trip records provided on the City of Chicago Data Portal\footnote{\url{https://data.cityofchicago.org}} and ii) the My Daily Travel household survey.

The City of Chicago Data Portal provides access to records of all ride-sourcing and taxi trips that transportation networking providers and taxi companies have reported to the City of Chicago for regulatory purposes since November 2018 and January 2013, respectively. The attributes of the trips are temporally and spatially aggregated to prevent a re-identification of individual trips. Each trip record includes information about the trip start and end times rounded to the nearest 15 minutes, the pick-up and drop-off community areas as well as the fare amount. The pick-up and drop-off census tracts are also available if at least three trips started or ended in the relevant census tract in the relevant 15-minute period. For ride-sourcing trips, it is also known if a pooled trip was requested. Thus, solo and pooled ride-sourcing trips can be distinguished.

The My Daily Travel household survey is a large-scale household travel survey that was conducted by the Chicago Metropolitan Agency for Planning (CMAP) between August 2018 and May 2019. The survey collected information about the daily travel behaviour of a representative sample of more than 12,000 households in North-eastern Illinois. More information about the survey is available in westat2020my. The collected data include trip records with information about the chosen transport mode (car, transit, bike, walking, taxi, solo-ride-sourcing or pooled ride-sourcing), trip start and end times as well as origin and destination census tracts.

For the construction of the estimation sample, we exploit the temporal and spatial overlap of the ride-sourcing and taxi trip records from the City of Chicago Data Portal and the My Daily Travel household survey. Consequently, we limit our analysis to trip records produced between 1 November 2018 and 3 May 2019. Since we are interested in understanding ride-sourcing use in the context of general travel demand, we only consider trips on weekdays between 5:00 and 23:00. Furthermore, to make it possible to impute the attributes of non-chosen alternatives, we restrict our analysis to trips for which the reported start and end locations are distinct. The start and end points of a trip are given by the centroids of the origin and destination census tracts or community areas. We exclude trips that start or end outside of the City of Chicago.

After applying these inclusion criteria, we are left with 12,593 trip records from the My Daily Travel household travel survey. For our analysis, we consider all trips from the My Daily Travel household survey that satisfy the inclusion criteria. 18,784,655 solo ride-sourcing, 7,290,921 pooled ride-sourcing and 3,821,709 taxi trips from the City of Chicago Data Portal satisfy the inclusion criteria.

We briefly describe the ride-sourcing and taxi trip records that meet the inclusion criteria. Figure (ref) shows the average weekday ride-sourcing and taxi trip counts by pick-up community area. It can be seen that the demand for ride-sourcing and taxi trips is concentrated in central zones (i.e. the Northeast) of the study area. In addition, Figure (ref) shows the average proportion of ride-sourcing trips requested as pooled trips by pick-up community area. We observe that the proportion of ride-sourcing trips requested as pooled trips is larger in the peripheral areas of the study area. Finally, Figure (ref) visualises the average ride-sourcing and taxi trip count in the whole study area by time of day. Whereas the demand for taxi is relatively balanced throughout the day, the demand for solo and pooled ride-sourcing trips exhibits pronounced morning and evening peaks.

figure[figure omitted — 310 chars of source]
figure[figure omitted — 233 chars of source]
figure[figure omitted — 204 chars of source]

Sampling protocol

If all eligible ride-sourcing and taxi trip records were considered, the resulting estimation sample would be too large to be processed. Therefore, we employ a choice-based sampling strategy to select ride-sourcing and taxi trips from the City of Chicago Data Portal. More specifically, we randomly select 20,000 records each from the sets of eligible solo ride-sourcing, pooled ride-sourcing and taxi trips. Table (ref) shows the absolute and relative frequencies of the observed mode choices in the two subsamples and the final estimation sample.

We use the My Daily Travel household survey to estimate the population mode share. In line with the sample design, the population quantity of interest is the mode share for trips between distinct census tracts in Chicago on weekdays between 5:00 and 23:00. To that end, we calculate average person trip rates of the population of Northeastern Illinois using the person-specific sampling weights provided in the My Daily Travel household survey. The calculated population mode share is shown in the column “Population---Share” in Table (ref).

table[table omitted — 1,185 chars of source]

Secondary data sources

After merging the two subsamples, we supplement the resulting dataset with information from secondary data sources.

The median household income and median age of each census tract are obtained from the American Community Survey uscb2021acs. The spatial distributions of the two quantities are visualised in Figure (ref).

We source various census tract attributes pertaining to employment and housing (employment density, residential density, employment and housing diversity), pedestrian friendliness (pedestrian network density, intersection density), transit supply (average proximity to transit, average transit service frequency) and car ownership (proportion of households with zero cars) from the Smart Location Database maintained by the US Environmental Protection Agency epa2021sld. Employment and housing diversity is an entropy-based diversity index accounting for employment numbers in five categories (retail, office, industrial, service and entertainment) and occupied housing from the database. Figures (ref)--(ref) show the spatial distributions of the extracted quantities.

Furthermore, information on seven land use categories (residential, commercial, institutional, industrial, transportation / communication / utilities / waste, agricultural, open space) is gathered from the 2013 CMAP Land Use Inventory cmap2015land. For each census tract in the study area, we calculate an entropy-based land use diversity index of the form $D = \left ( \sum_{c \in \mathcal{C}} p_c \ln p_c \right ) \big/ \vert \mathcal{C} \vert$, where $\mathcal{C}$ denotes the set of considered land use categories, and $p_c$ is the proportion of land use classified as category $c$. The right panel of Figure (ref) shows the spatial distribution of the calculated diversity index.

In addition, information about park fees are sourced from the CMAP Parking Inventory ghaffar2020modeling. The spatial distribution of the average hourly park rate is shown in the right panel of Figure (ref).

Finally, weather information is taken from daily meteorological summaries for O'Hare International Airport provided by ncei2021daily. The average daily temperature and total daily precipitation during the observation period are shown in Figure (ref).

figure[figure omitted — 256 chars of source]
figure[figure omitted — 244 chars of source]
figure[figure omitted — 277 chars of source]
figure[figure omitted — 272 chars of source]
figure[figure omitted — 289 chars of source]
figure[figure omitted — 281 chars of source]
figure[figure omitted — 236 chars of source]

Imputation of mode attributes

Lastly, we impute the attributes of the mode choice alternatives. Driving times and distances, transit connections as well as walking and bicycling travel times are obtained from the HERE Routing Application Programming Interface\footnote{\url{https://developer.here.com/products/routing}}.

For the calculation of the cost of the driving alternative, we consider two variable cost components, a vehicle running cost component of 0.20 \$/mile and a fuel cost component. To compute the latter, we assume a fuel economy of 20 miles per gallon. Weekly average retail gasoline prices are sourced from eia2021chicago. Figure (ref) shows the evolution of the unit price of gasoline during the observation period. We approximate transit fares using agency-specific revenue information provided in the 2019 National Transit Database fta2021ntd. Based on this information, we assume a fare of 0.50 \$/mile for bus and a fare of 0.30 \$/mile for rail and metro.

Since passenger wait times for ride-sourcing and taxi are not observed, we assume a fixed wait time of two minutes for solo ride-sourcing, pooled ride-sourcing and taxi. These waiting times are included in the total travel times of these alternatives. To account for possible detours for picking up other passenger during pooled ride-sourcing trips, we add a ten percent travel time penalty to the driving time of pooled ride-sourcing.

We use random forests to impute solo and pooled ride-sourcing fares. The random forest models take into account lagged fare information (i.e. the 25\textsuperscript{th}, 50\textsuperscript{th} and 95\textsuperscript{th} percentiles of the per kilometre fare in the whole network during the 30-minute time period preceding the start time of the trip), trip attributes (driving distance and time, start time, the day of the week), atmospheric conditions (average daily temperature, precipitation) as well as various attributes of the origin and destination census tracts. Taxi fares are calculated using official fare information chicago2020taxi.

Table (ref) provides a summary of the attributes of the chosen alternatives.

figure[figure omitted — 161 chars of source]
table[table omitted — 1,927 chars of source]

Econometric model

Constructing the estimation sample from two revealed preference data sources creates two challenges in the development of a discrete choice model. First, a sampling correction is needed to account for the enrichment of the household travel survey data with ride-sourcing and taxi trip records. Second, the constructed revealed preference mode choice dataset is likely to exhibit endogeneity biases, because the demand for the ride-sourcing options and the price of the ride-sourcing options are simultaneously influenced by supply-side constraints and surge pricing mechanisms. In this section, we present a MEV-based discrete choice model with sampling and endogeneity corrections to address these challenges. First, we develop the sampling correction (Section (ref)). Here, we consider a conditional maximum likelihood estimator due to its superior efficiency properties. Then, we describe the endogeneity correction (Section (ref)). Here, we select the control function approach due to its simplicity. The reader is directed bierlaire2020sampling for a recent review of sampling correction approaches in discrete choice analysis and to mcfadden1999chapter for an earlier synthesis of the topic. guevara2015critical provides a review of endogeneity correction approaches.

Discrete choice analysis under non-random sampling

We consider a sample of $N$ individuals indexed by $n = 1, \ldots, N$. Every individual $n$ is observed to choose an alternative $y_n$ from the set $\mathcal{M} = \{1, \ldots, J \}$. We stipulate a parametric model which generates the probability that individual $n$ chooses alternative $j \in \mathcal{M}$ given explanatory variables $\boldsymbol{x}_n$ with density $\mu(\boldsymbol{x}_n)$ and the unknown parameter $\boldsymbol{\theta}$:

equation[equation omitted — 90 chars of source]

Random utility theory mcfadden1981econometric posits that a rational decision-maker $n$ chooses the option $y_n$ with the highest utility from $\mathcal{M}$, i.e.

equation[equation omitted — 83 chars of source]

whereby

equation[equation omitted — 92 chars of source]

denotes the utility of alternative $j \in \mathcal{M}$. The utility $U_{nj}$ is decomposed into a deterministic aspect $V_{nj}(\boldsymbol{x}_{nj}, \boldsymbol{\theta})$ and a stochastic aspect $\varepsilon_{nj}$, which is unknown to the analyst.

We further suppose that the sample consists of $S$ subsamples indexed by $s = 1, \ldots, S$. Each subsample $s$ is characterised by a sampling protocol involving endogenous and exogenous stratification. Under an exogenous sampling protocol, the analyst selects observations based on exogenous variables $\boldsymbol{x}$.\footnote{Note that the exogenous variables $\boldsymbol{x}$ in the exogenous sampling protocol need not be the same as the explanatory variables $\boldsymbol{x}$ in the choice model.} Under an endogenous sampling protocol, the analyst selects observations based on realised choices.

To develop a choice model considering both exogenous and endogenous stratification, we let

equation[equation omitted — 74 chars of source]

denote the probability that a population member with configuration $\{j, \boldsymbol{x}_n \}$ qualifies for the subpopulation from which subsample $s$ is drawn. Consequently, the joint probability of observing a case with configuration $\{j, \boldsymbol{x}_n \}$ that qualifies for subsample $s$ is $R_s (j, \boldsymbol{x}_n) P(j \vert \boldsymbol{x}_n ; \boldsymbol{\theta}) \mu(\boldsymbol{x}_n)$, and the following marginal probability gives the population share $Q_s$ of the subpopulation from which subsample $s$ is recruited:

equation[equation omitted — 134 chars of source]

Then, by Bayes' rule, the probability of observing a case with configuration $\{j, \boldsymbol{x}_n \}$ conditional on membership in subsample $s$ is given by

equation[equation omitted — 201 chars of source]

Now, the likelihood of configuration $\{j, \boldsymbol{x}_n \}$ unconditional on subsample membership is

equation[equation omitted — 138 chars of source]

where $H_s = \frac{N_s}{N}$ is the share of subsample $s$ in the total sample. Hence, the likelihood of a case with choice $j$ given explanatory variables $\boldsymbol{x}_n$ is

equation[equation omitted — 379 chars of source]

Substituting ((ref)) for $P(j, \boldsymbol{x}_n \vert s; \boldsymbol{\theta})$ yields a likelihood which is independent of $\mu (\boldsymbol{x}_n)$. We have

equation[equation omitted — 305 chars of source]

which simplifies to

equation[equation omitted — 249 chars of source]

with

equation[equation omitted — 87 chars of source]

((ref)) suggests a conditional maximum likelihood estimator manski1981alternative of the form

equation[equation omitted — 293 chars of source]

whereby the term

equation[equation omitted — 202 chars of source]

needs to be adapted to the stipulated parametric form of the choice model ((ref)).

In the MEV family of discrete choice models mcfadden1978modeling, the probability of choosing alternative $j \in \mathcal{M}$ conditional on explanatory variables $\boldsymbol{x}_n$ and parameters $\boldsymbol{\theta}$ is given by

equation[equation omitted — 209 chars of source]

where

equation[equation omitted — 250 chars of source]

with

equation[equation omitted — 76 chars of source]

and

equation[equation omitted — 168 chars of source]

Here, $V_{nj'}$ is the deterministic aspect of utility, which depends on explanatory variables $\boldsymbol{x}_n$ and parameter $\boldsymbol{\beta}$. $G(\psi_{n1}, \ldots, \psi_{nJ}; \boldsymbol{\lambda})$ is a MEV generating function with parameter $\boldsymbol{\lambda}$.

To evaluate ((ref)) under the MEV assumption, we generalise the result presented in bierlaire2008estimation from purely choice-based samples to a wider class of enriched samples. We have

equation[equation omitted — 240 chars of source]

and define

equation[equation omitted — 205 chars of source]

Consequently, we obtain

equation[equation omitted — 510 chars of source]

Thus, a conditional maximum likelihood estimator for the MEV family of discrete choice models under non-random sampling is given by

equation[equation omitted — 438 chars of source]

with $\boldsymbol{\theta} = \{ \boldsymbol{\beta}, \boldsymbol{\lambda} \}$.

Control function correction of endogeneity

We partition the explanatory variables $\boldsymbol{x}$ into exogenous explanatory variables $\boldsymbol{c}$ and an endogenous explanatory variable $p$ such that $\boldsymbol{x} = \{\boldsymbol{c}, p \}$. Then, the utility of alternative $j \in \mathcal{M}$ is

equation[equation omitted — 124 chars of source]

with

equation[equation omitted — 132 chars of source]

Here, $\boldsymbol{z}_{nj}$ denotes a set of instruments, $\boldsymbol{\gamma}$ and $\boldsymbol{\delta}$ are unknown parameters, and $\xi_{nj}$ is an error term. The error term $\xi_{nj}$ captures the influence of unobserved attributes of alternative $j$ which impact $p_{nj}$ but are not included in $\boldsymbol{z}_{nj}$ and $\boldsymbol{c}_{nj}$. The instruments $\boldsymbol{z}_{nj}$ and exogenous explanatory variables $\boldsymbol{c}_{nj}$ are independent of the stochastic aspect of utility $\varepsilon_{nj}$ and the stochastic disturbance $\xi_{nj}$. Yet, the endogenous explanatory variable $p_{nj}$ is correlated with $\varepsilon_{nj}$, i.e. $\text{Cov}(p_{nj}, \varepsilon_{nj}) \neq 0$ and thus $\text{Cov}(\xi_{nj}, \varepsilon_{nj}) \neq 0$. Ignoring the endogeneity of $p_{nj}$ in the estimation of the choice model parameters $\boldsymbol{\theta}$ leads to inconsistent parameter estimates train2009discrete.

The control function correction of endogeneity petrin2010control consists of constructing a control variable which, when included into the utility specification, absorbs the aspect of $\varepsilon_{nj}$ that is correlated with $p_{nj}$. The utility error is decomposed as

equation[equation omitted — 101 chars of source]

where $C(\xi_{nj}, \phi)$ is the control function with parameter $\phi$. $\tilde{\varepsilon}_{nj}$ is the residual error, which remains after conditioning out the aspect of $\varepsilon_{nj}$ that is correlated with $p_{nj}$. The simplest specification of the control function is

equation[equation omitted — 50 chars of source]

whereby $\phi$ is an unknown scalar parameter. Then, the utility ((ref)) writes

equation[equation omitted — 119 chars of source]

A choice model with a control function correction of endogeneity is estimated in two stages. First, the endogenous variable $\boldsymbol{p}_{j}$ is regressed on the instruments $\boldsymbol{z}_{j}$ and exogenous variables $\boldsymbol{c}_{nj}$. The residuals $\widehat{\xi}_{nj}$ from this regression are used to calculate the control function. In the second stage, the choice model is estimated, with the control function being included in the utility equation.

Model specification

Second stage: Discrete choice model

In our analysis of the mode choice dataset described in Section (ref), the second stage of the two-stage model introduced in the previous section is specified as a multinomial logit model. We also explored various nested and cross-nested logit model specifications using the conditional maximum likelihood estimator exhibited in ((ref)), but no meaningful nesting structures emerged.

The multinomial logit model assumes a specification of the random utility of the following form: We let

align[align omitted — 340 chars of source]

with

equation[equation omitted — 232 chars of source]

Here, $\boldsymbol{x}_{nj}^{\text{(alt.-spec.)}}$ and $\boldsymbol{x}_{nj}^{\text{(trip-spec.)}}$ denote alternative- and trip-specific attributes, respectively. The corresponding parameters are denoted by $\boldsymbol{\beta}^{\text{(alt.-spec.)}}$ and $\boldsymbol{\beta}_{j}^{\text{(trip-spec.)}}$, respectively. Whereas alternative-specific attributes vary across alternatives and trips (e.g. travel time, travel cost etc.), trip-specific attributes only vary across trips (e.g. census tract attributes at the origin and destination). The parameters $\boldsymbol{\beta}_{j}^{\text{(trip-spec.)}}$ pertaining to trip-specific attributes are necessarily alternative-specific. For identification, we fix $\boldsymbol{\beta}_{\text{car}}^{\text{(trip-spec.)}}$ to zero.

We also incorporate alternative-specific departure time preferences in the utility specification. Following earlier studies on air-travel itinerary choices koppelman2008schedule, lurkin2017accounting, wen2020incorporating, we consider continuous representations of departure time preferences using a weighted sum of sine and cosine functions. The specification has the following form:

equation[equation omitted — 430 chars of source]

Here, $\beta_{j,1}, \ldots, \beta_{j,6}$ are unknown parameters. $t_{n}$ is the observed departure time of trip $n$ in minutes past midnight, 1440 is the total number of minutes in a day. Compared to discrete representations, continuous representations of departure time preferences produce more realistic demand predictions due to their smoothing properties koppelman2008schedule, lurkin2017accounting.

In accordance with the model formulation put forward in the previous section, $\phi_{j}\widehat{\xi}_{nj}$ in ((ref)) is the control function, and $\phi_{j}$ is the corresponding coefficient. $\varepsilon_{nj}, \tilde{\varepsilon}_{nj}$ are error terms, which are assumed to be independent and identically distributed according to $\text{Gumbel}(0,1)$ across $n, t$. $\tilde{\varepsilon}_{nj}$ is the residual utility error that remains after conditioning on the aspect of $\varepsilon_{nj}$ that is correlated with the endogenous variable.

First stage: Control function

We hypothesise that the prices of solo and pooled ride-sourcing are endogenous, because the demand for the ride-sourcing options and their prices are co-determined by supply-side constraints and surge pricing mechanisms. We employ the control function approach described in Section (ref) to correct for this price endogeneity. To form the control function, we must find suitable instruments that are i) correlated with the endogenous variable (i.e. price) and ii) not correlated with the error term of the demand equation. nevo2000practitioner distinguishes three types of demand-side instrument, namely i) cost shifters, ii) non-price attributes of other alternatives---also referred to as BLP-type instruments berry1995automobile---and iii) prices of the same alternative in other markets---also referred to as Hausman-type instruments hausman1994competitive, hausman1996valuation.

In this work, we consider cost-shifters and non-price attributes of other alternatives as demand-side instruments for the prices of solo and pooled ride-sourcing. First, we include driving cost as a cost-shifting instruments, whereby driving cost is calculated as the driving distance in miles times the retail price of gasoline in USD per gallon divided by an assumed fuel economy of 20 miles per gallon. Second, we consider the aggregate frequency of transit service per capita in the origin census tract as BLP-type instruments.

Estimation practicalities

The multinomial logit models are estimated using the conditional maximum likelihood estimator given in ((ref)). Note that the presented estimator is fully general in that it can account for both endogenous and exogenous stratification. In the current application, the sampling protocol is purely choice-based and does not involve stratification by an exogenous variable. Thus, we have $\ln \alpha_{nj} = \ln \alpha_{j} \; \forall \; n \in \{1, \ldots, N \}$. $\alpha_{j}$ is given by the sampling protocol defined in Table (ref). Specifically, we have $\alpha_{j} = H_{j} / Q_{j}$, whereby $H_{j}$ is the share of observations in the sample choosing alternative $j$ (see column “Estimation sample---Share” in Table (ref)), and $Q_{j}$ is the share of the population choosing alternative $j$ (see column “Population---Share” in Table (ref)).

We implement the conditional maximum likelihood estimator for the multinomial logit models using PandasBiogeme bierlaire2018pandasbiogeme. The standard errors of the parameters of the discrete choice model with a control function correction are bootstrapped using 100 resamples. The first-stage regressions of the two-stage model are estimated using ordinary least squares.

Results

In Table (ref), we provide the estimation results of an uncorrected model (without control function correction of endogeneity but with sampling correction) and a corrected model (with endogeneity and sampling corrections). The first-stage results of the two-stage model are presented in Table (ref). Summary statistics for the first-stage regressions, namely $F$-statistics, the associated $p$-values and the coefficients of determination $R^2$ are given in Table (ref). Summary statistics for the second-stage models are given in Table (ref).

First, we test for the presence of endogeneity. Under the null hypothesis that the prices of the ride-sourcing alternatives are exogenous, the second-stage coefficients $\phi_{\text{solo ride-sourcing}}$ and $\phi_{\text{pooled ride-sourcing}}$ on the first-stage residuals are zero. Note that the considered uncorrected model is nested within the considered corrected model, since the uncorrected model can be obtained from the corrected model by setting $\phi_{\text{solo ride-sourcing}}$ and $\phi_{\text{pooled ride-sourcing}}$ equal to zero. Under the null hypothesis that the prices of the ride-sourcing alternatives are exogenous, the restrictions imposed by the uncorrected model are supported by the observed data. Table (ref) shows that the log-likelihood of the uncorrected model is $-75478.51$, while the log-likelihood of the corrected model is $-75405.76$. A log-likelihood ratio test indicates that the improvement in fit offered by the corrected model is statistically significant ($\tilde{\chi}^{2} = 145.51$, $\text{df} = 2$, $p < 0.001$). Thus, we reject the constraints of the uncorrected model and conclude that the prices of the ride-sourcing alternatives are endogenous.

In the corrected model, the estimates of the parameters pertaining to mode attributes have the expected signs and are significantly different from zero. More precisely, the mode-specific travel time parameters are all negative and statistically significant. As expected, a larger number of transfers appears to decrease the propensity of choosing public transit, and a higher hourly park rate at the destination appears to decrease the propensity of choosing car. Strikingly, the estimate of $\beta_{\text{cost, car}}$ is not statistically significant in the uncorrected model.

In Table (ref), we compare the weighted direct aggregate arc elasticities with respect to travel cost and time of the two models. It can be seen that in the corrected model, the demand for taxi as well as for solo and pooled ride-sourcing is substantially more elastic with respect to price than in the uncorrected model. For example, the estimated direct aggregate arc elasticity with respect to the cost of pooled ride-sourcing is $-0.262$ in the uncorrected model and $-0.923$ in the corrected model. The corrected model further reveals that the demand for the considered mode choice alternatives is elastic with respect to travel time. As expected, walking and biking exhibit the highest elasticities with respect to travel time. Taxi and pooled ride-sourcing are more elastic with respect to travel time than solo ride-sourcing.

Figure (ref) visualises the estimated continuous departure time preferences in the corrected model. While there are minor differences in departure time preferences across the three modes in the morning, afternoon and evening hours, solo and pooled ride-sourcing appear to be comparatively less likely to be chosen mid-day.

We also observe that weather conditions affect the demand for taxi and ride-sourcing. Our utility specification includes both main and interaction effects of the average daily temperature and the daily precipitation amount. To facilitate the interpretation of the effects, we standardised the former and kept the latter on its original scale. The estimates of these effects are statistically significant for taxi as well as solo and pooled ride-sourcing. Since the estimated effects have the same signs and are in the same order of magnitude, the same interpretation applies to the estimated effects for all three modes. At the mean average daily temperature in the observation period, positive precipitation increases the demand for taxi and ride-sourcing. On dry days, a higher temperature leads to increased demand for taxi and ride-sourcing.

The corrected model also provides insights into the influence of census tract attributes on travel demand. Due to the inclusion of quadratic terms in the utility specification, we are able to capture non-linear income and age effects on the demand for taxi as well as solo and pooled ride-sourcing. These effects are visualised in Figure (ref) over their respective realised ranges in the training dataset. For all three modes, the non-linear income effects are concave down, whereby the curvature is more pronounced for taxi and solo ride-sourcing than for pooled ride-sourcing. The demand for pooled ride-sourcing appears to be less sensitive to income compared to the demand for taxi and solo ride-sourcing. Increasing income initially has a positive effect on the demand for taxi and solo ride-sourcing, but the effect of income becomes negative for median annual household incomes above USD 140,000. This suggests that the demand for taxi and solo ride-sourcing is comparatively lower in census tracts with high household incomes.

Next, we consider the estimated age effects in the corrected model. For solo and pooled ride-sourcing, the age effects are concave down, while they are concave up for taxi. The curvature of the age effect on demand for taxi is substantially more pronounced than for the other two modes. In comparison to the age effect for taxi, solo and pooled ride-sourcing do not appear sensitive to age. The effect of age on taxi demand increases sharply for median ages above 35 years, which suggests that taxi demand is comparatively higher in census tracts with older residents.

Various land use and built environment characteristics also influence the demand demand for taxi and ride-sourcing. For example, a higher residential density increases the propensities of choosing taxi and solo ride-sourcing. A higher employment density increases the propensity of choosing taxi but decreases the propensity of choosing ride-sourcing. A higher land use diversity decreases the propensities of choosing taxi and ride-sourcing. A denser network of pedestrian-oriented links decreases the propensities of choosing taxi and ride-sourcing. However, a higher intersection density at the trip origin increases the propensities of choosing taxi and ride-sourcing.

table[table omitted — 11,961 chars of source]
table[table omitted — 795 chars of source]
table[table omitted — 372 chars of source]
table[table omitted — 339 chars of source]
table[table omitted — 632 chars of source]
figure[figure omitted — 133 chars of source]
figure[figure omitted — 237 chars of source]

Welfare analysis

We also use the corrected model to analyse the welfare implications of ride-sourcing. More specifically, we consider three scenarios in which we simulate welfare losses due to the removal of i) all ride-sourcing, ii) solo ride-sourcing and iii) pooled ride-sourcing services from the choice sets of observations in which the removed mode is the chosen mode. In addition, we analyse the welfare implications of ride-sourcing taxes, inspired by a congestion tax implemented in Chicago in 2020 mcmahon2020citys. More specifically, we consider two taxation scenarios in which a tax is added to the fare of trips in which solo ride-sourcing is selected. In the first scenario, we impose a fixed tax of USD 3 on solo ride-sourcing trips, and in the second scenario, we apply a variable tax of 20% added to solo ride-sourcing fares.

For each scenario and trip, we compute compensating variations, i.e.\ the monetary compensations that offset the alteration of the choice sets. Since the considered model includes alternative-specific cost parameters, it is not possible to analytically compute compensating variations. Therefore, we adopt the simulation approach presented in mcfadden2012computing. For completeness, we also describe the approach in Appendix (ref).

In Figures (ref) and (ref), we show box plots of the computed compensating variations in the elimination and taxation scenarios, respectively. Figure (ref) suggests that welfare losses are largest due to the elimination of all ride-sourcing services, closely followed by the removal of only solo ride-sourcing services, whereas welfare losses due to the removal of pooled ride-sourcing services are comparatively small. While the mean compensating variations for the first two elimination scenarios are USD 0.44 and USD 0.35, respectively, the compensating variation in the third scenario, in which only pooled ride-sourcing services are eliminated, is USD 0.13. As expected, welfare losses in the taxation scenarios are smaller compared to the elimination scenarios (see Figure (ref)). This is because in the taxation scenarios alternatives are made less attractive through the introduction of a tax but are not entirely removed from choice sets. It can be seen that welfare losses are higher in the scenarios with a fixed tax compared to the scenarios with a variable tax. Whereas the mean compensating variation is USD 0.09 in the first taxation scenario, the mean compensating variation in the second scenario is USD 0.05. In all scenarios, the distributions of the compensating variations exhibit a considerable spread and are right-tailed. For example, in the first elimination scenario, the interquartile range of the compensating variations is USD 0.41, and in the first taxation scenario the interquartile range is USD 0.12. In all scenarios, the mean is larger than the median. These results suggest that the distribution of ride-sourcing benefits is highly heterogeneous. Overall, the compensating variations appear small. However, this can be explained by the fact that the ride-sourcing alternatives have comparatively small probabilities of being selected compared to other alternatives, such as car and public transit.

In Figure (ref), we present the average compensating variations by community area for the elimination scenarios, in which services are removed from choice sets. The figure reveals substantial heterogeneity in the distributions of the computed compensating variations in the three considered scenarios across the community areas of the study region. In all three scenarios, ride-sourcing benefits are valued higher in central areas. The average compensating variations in the community areas of the study area in the first scenario range from USD 0.09 to 0.70. Benefits of solo ride-sourcing are valued higher than the benefits of pooled ride-sourcing. Whereas the average compensating variations in the second scenario range from USD 0.03 to 0.55, the average compensating variations in the third scenario range from USD 0.04 to only 0.31.

Finally, in Figure (ref), we present the average compensating variations by community area for the taxation scenarios. The average compensating variations range from USD 0.01 to 0.13 in the fixed tax scenario and from USD 0.01 to 0.0.08 in the variable tax scenarios. The spatial distributions of the compensating variations appear similar in both scenarios and are consistent with the spatial distributions obtained in the elimination scenarios.

figure[figure omitted — 256 chars of source]
figure[figure omitted — 249 chars of source]
figure[figure omitted — 355 chars of source]
figure[figure omitted — 285 chars of source]

Conclusion

In this paper, we presented and applied an approach for estimating ride-sourcing demand at a disaggregate level from multiple data sources using DCMs. In sum, our research makes four contributions to the literature. First, we demonstrate how ride-sourcing demand estimation with DCMs can be performed by fusing multiple disaggregate data sources. Second, we show how traditional household travel surveys can be enriched with emerging sources of big data (i.e. trip records). Third, we highlight the importance of controlling for endogeneity biases in ride-sourcing demand estimation. Finally, we provide a methodology for incorporating emerging mobility options (such as ride- and bike-sharing etc.) into disaggregate activity-based travel demand forecasting models.

There are several ways in which our work could be extended. First, an integrated choice and latent variable model could be adopted to accommodate flexible substitution patterns and to simultaneously estimate the two model stages. Second, the constructed mode choice dataset could be enriched with trip records providing information about other emerging transport modes such as bike-sharing. Third, the temporal structure of the data could be explicitly considered to investigate the temporal stability of the structural relationship between travel demand and the various explanatory variables.

Author contribution statement

Rico Krueger: Conceptualisation, Methodology, Software, Formal analysis, Investigation, Data curation, Writing – original draft, Visualisation. Michel Bierlaire: Conceptualisation, Methodology, Writing – review & editing. Prateek Bansal: Conceptualisation, Methodology, Investigation, Writing – original draft.