Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
107,246 characters · 10 sections · 71 citation commands
The learning effects of subsidies to bundled goods: a semiparametric approach $^*$
\thispagestyle{plain}
In markets for innovations, it is reasonable to assume that consumers may be uncertain about product quality. In these settings, firms may offer a bundle, which combines an innovation with other goods of better known quality, in order to mitigate consumer uncertainty and induce learning with respect to the novelty Reinders2010. Typical examples include packaging trial versions of new software with updates of better known software Sheng2009, and the bundling of financial innovations with consolidated services tufano2008using.
In these innovation settings, bundles are often offered at a discount. Such subsidy is expected to increase learning about the innovation. Indeed, inasmuch as substitution effects dominate, we may expect users to increase their demand of the bundled good in response to the subsidy, thereby acquiring additional information about the innovation. Moreover, a forward-looking agent may increase demand due to an additional experimentation or learning motive, i.e. in order to increase the informational bequest to be left for her future selves.
Do temporary subsidies to bundled goods induce long-term changes in consumption behaviour due to increased learning on the relative quality of one of the constituent goods? This paper provides theoretical guidance and empirical evidence on the role of this mechanism. Theoretically, we introduce a dynamic model where an agent learns about the quality of an innovation on an essential good through repeated consumption. The agent has a target level of consumption of an essential good, as well as preferences over consumption in other goods. The target can be fulfilled by relying on a known outside option, an innovation whose relative quality is uncertain, or a bundle with a share $\phi$ of the outside option and $1-\phi$ of the unknown good. Increasing consumption of either the bundle or the innovation increases signal precision regarding the underlying quality, albeit at possibly different rates. In this setting, we show that the effect of a one-period subsidy on the bundle can be decomposed into a direct price effect, and indirect effects due to a learning or experimentation motive, as the agent leverages the discount to increase the informational bequest to be left to her future selves. Our results further show that subsidies may affect not only the mean, but also the dispersion of demand, and that, given the essentiality nature of the innovation, any extemporaneous effects of a one-period subsidy to the bundle are fully attributable to learning.
Empirically, we assess the predictions of our theory of learning in the context of a randomised experiment conducted in partnership with a large ridesharing service in S\ {a}o Paulo, Brazil. The experiment randomly provided two ten-day discount regimes to a subset of the platform's user base. One of the regimes provided two daily 20% discounts, limited to 10 BRL per trip (around 4 USD at the time), on rides starting or ending at a train/metro station (integrated rides). The other regime provided two daily 50% discounts, also limited to 10 BRL per trip, on integrated rides. We also have access to a control arm, which was not informed about the experiment. For each individual in the experimental sample, we observe baseline survey answers on socioeconomic information and commuting habits, as well as the user history of demand for integrated and non-integrated (total minus integrated) rides on a fortnightly basis, starting four months prior to the experiment and going over six months after the discount ends. We view this setting as an ideal environment to test the predictions of our theory, as we observe a one-off unanticipated discount to a bundled good (integrated rides), which may generate information about the underlying quality of an innovation (non-integrated rides)\footnote{Ridesharing apps were authorised to operate in São Paulo in mid 2016; and the experiment was conducted in late 2018.} on an essential good (transportation services). We are also able to assess long-term effects of discounts, given the time range spanning over six months after the experiment took place.
The distribution of individual demand for rides exhibits heavy tails, which leads to inferences based on difference-in-means estimators performing poorly Lewis2015. Following Athey2023, we propose a semiparametric model for potential outcomes that, while constraining treatment effect heterogeneity, enables the construction of more efficient estimators. Motivated by our theory, we consider a parametrisation that allows treatment effects to change both the location and the scale of the demand for rides. By stratifying our analysis, we also allow for heterogeneity according to user's information sets. We then introduce a semiparametric estimator for our proposed method by relying on L-moments, a robust alternative to standard moments that characterise any distribution with finite mean Hosking1990. Our estimator, a semiparametric version of the generalized method of L-moments estimator of alvarez2023inference, has several attractive properties: in our adopted parametrisation, it can be easily computed as a generalised least squares estimator in a modified dataset; it yields as an immediate byproduct a specification (overidentifying restrictions) test of the adopted model; and, in the setting of Athey2023, it is asymptotically efficient without requiring any additional corrections, such as estimation of the model's efficient influence function Klaasen1987,Newey1990.
When applied to our experimental data, the semiparametric methods shed light on several patterns. The 50% discount leads to a large contemporaneous increase in integrated rides, amounting to 60% of the control group average at the period. However, there do not appear to be subsequent effects in either the mean or dispersion of integrated rides. As for the demand for non-integrated rides, we observe a sizeable and persistent decrease in their mean and dispersion in the fortnights after the discount took place. These effects last for over four months after the end of the experiment. Reductions in average non-integrated rides amount to 15% of the control group mean three months after the discount. We observe similar, albeit weaker, effects on the 20% discount regime, though in this case the contemporaneous change in average integrated rides is not statistically significant.
As we argue later on, observed patterns are mostly consistent with our learning model. Indeed, we may interpret our results through the following mechanism of our theory: upon increasing consumption as a response to the discount, agents increase learning about the relative (to the outside option) quality of ridesharing apps, updating their beliefs to a more pessimistic level. This reduces the demand for non-integrated rides. However, given the essentiality nature of transportation services, the consumer must shift consumption to other modes of transportation in order to meet her target. This substitution effect approximately compensates the impact of the pessimistic update on integrated rides -- recall the effect of quality on the bundle is dampened by the share $\phi$ --, and we do not observe changes in the demand for this type of trip after the discount is over. Note that, as a consequence of this mechanism, we would also expect an increase in the demand for the (unobserved) outside option in subsequent periods. Finally, to shed more light on the learning mechanism, we combine our experimental estimates with survey data to calibrate some key parameters of our model of learning. We find that, for the 50% discount, around 40% to 50% of the contemporaneous increase in integrated rides may be attributed to the experimentation motive, i.e. to a desire to accumulate more information on transportation services for future use.
Our paper is connected to a larger body of research that seeks to provide a rationale for the existence of bundled goods. Typical explanations range from cost-saving and preference-complementarity arguments Pyndick, to the use of bundles as a means of implementing price discrimination Adams1976, as well as a mechanism used by incumbent firms to deter market entry carlton2002strategic,nalebuff2004bundling. Firms may also offer bundles as a means of extracting surplus when product search is costly to consumers Harris2006, and firms have access to a technology that is able to predict those goods most suitable to a consumer, e.g. recommender systems in streaming services hiller4415301digital. In this paper, we provide theoretical as well as empirical evidence on yet another role played by bundles: they may be used to mitigate consumer uncertainty regarding an innovation and induce long-run changes in consumer behaviour via learning effects. Furthermore, by providing a formal theoretical foundation on the learning effects of subsidies to bundles and in-the-field experimental evidence on the role of this mechanism, we contribute to a body of research in marketing and product development that previously relied on laboratory and observational data to discuss bundling strategies for new product introduction SIMONIN1995219,Sheng2009,Reinders2010.
The present work is also related to the literature on the markets and pricing policies of experimentation goods shapiro1983optimal,liebeskind1989markets,bergemann2006dynamic,Chen2022. Experimentation goods are commodities whose underlying quality is only (partially) revealed after consumption. Our paper provides evidence that bundling, when coupled with a temporary subsidy, may be a convenient strategy to induce long-run changes in demand for this class of goods.\footnote{In a two-period setting, shapiro1983optimal shows that, if consumers initial expectations regarding the quality of the experimentation good are pessimistic, the optimal strategy of a monopolist would be to provide an initial subsidy to increase adoption. In such situation, our results suggest that bundles may offer a complementary, and possibly cheaper, approach to induce these desired changes.} The nature of our experimental variation is such that our results follow without any assumptions on market equilibrium or firm behaviour. As a consequence, we view our paper as documenting a mechanism in consumer behaviour that may be incorporated in future research on the optimal design of pricing policies in these settings.
From the econometric standpoint, our paper contributes to the literature on semiparametric estimation Newey1990,Bickel1993,Newey1995,Bolthausen2002,kosorok2008introduction. We introduce an alternative estimator for Athey2023's model that is asymptotically efficient without requiring any further adjustments, e.g. the correction by a cross-fitted estimate of the model efficient influence function employed by Athey2023. Our estimator is computationally convenient; and its objective function immediately provides us with a J-test of over-identifying restrictions. More generally, we view our paper as introducing the generalized method of L-moments approach of alvarez2023inference -- which, in a parametric setting, improves upon maximum likelihood estimation in finite samples of popular distributions whilst remaining asymptotically equivalent to it --, to a semiparametric environment.\footnote{alvarez2023inference consider an extension of their main approach that can be used to fit a parametric model for the error term of a semiparametric model that is estimated in a first step. Their proposed construction and the range of applications covered by their method are quite distinct from ours, though. In particular, their method requires estimating an influence-function correction to account for first-step estimation, whereas our method does not rely on knowledge or estimation of any influence function.} Given the attractive statistical properties of L-moments (see (ref) and Online (ref) for a discussion; also Alvarez2023_mixture, Alvarez2023_mixture), we view this paper as a further step in developing generalised method of L-moment estimation in a semiparametric setting.\footnote{In a recent paper, Alvarez2023_mixture show that a sieve-like extension of the estimator in alvarez2023inference provides a computationally convenient estimator with valid inferential guarantees in nonparametric quantile mixture models.}
Finally, our paper is also connected to the literature on the Economics of Urban Transportation. Methodologically, we introduce a dynamic model of demand for transportation services, which contrasts with the discrete choice models typically adopted in the literature small2007economics. Our modelling approach may be more broadly useful to generate predictions in settings where learning is an important component, and individual choices are observed as aggregates over repeated time windows. Our paper also contributes to the literature on behavioural interventions in transportation markets METCALFE2012503,Kristal2019,GRAVERT20211. In both developing and developed countries, a long-run declining trend in the use of public trasnportation services has been documented mallett2018trends,Rabay2021. This decline has been further accentuated by the Covid-19 emergency mallett2022public,loh2023ensuring, with a full recovery to pre-pandemic patterns very much unlikely absent further intervention Dai2021,Tsavadari. Given the high fixed costs of operating a public transportation system, these changes tend to threaten the long-run financial viability of these services, which often translates into increased government subsidies flowing into the operation welle2020safer,aguilar2021after,Tsavadari. Insofar as the outside option in our analysis may be interpreted as public transportation,\footnote{As we argue in more detail in (ref), we restrict our empirical analysis to a subsample of users for which the outside option is more likely to be interpreted as the train/metro system. Moreover, we note that, in our subsample of interest, public transportation usage is quite prevalent, with 68.8% of participants reporting the train, metro or public bus system as one of their main modes of transportation.} our results would indicate that temporary subsidies to modal integration may induce persistently higher public transportation takeup by improving beliefs regarding the relative quality of the service.\footnote{Public transportation trips would increase because the average number of integrated rides does not decrease and, when interpreting the results of our experiment through the lenses of our theory, we expect an increase in the outside option after the discount is over.} This result suggests that these types of policies could be employed to partially offset long-run trends in commuting behaviour.
The remainder of this paper is organised as follows. Section (ref) introduces our model. Section (ref) discusses the experimental design. Section (ref) presents our target estimands and the proposed semiparametric model. Section (ref) discusses estimation. Section (ref) presents the results of our empirical analysis. Section (ref) concludes. The \hyperref[supplementary]{Online Appendix} contains the proofs of our main results, as well as additional details on the semiparametric estimator.
In this section, we introduce a model which aims to capture the main mechanisms associated with discounts to bundled goods in a setting where the quality of one of the elements in the bundle is unknown. We focus on an environment where the goods of unknown quality are essential, in the sense that the consumer has a target level of consumption required in order to be able to consume other goods. Our leading example are transportation services, which are the focus of our empirical application. It seems reasonable to assume that preferences for commuting services arise due to having to work or study outside home, which generates a target level of transportation needs that must be met.\footnote{Indeed, this approach is in keeping with both textbook discussions on the demand for transportation services small2007economics, as well as more recent treatments Kreindler2023.}
Consider a consumer which, at every period $t \in \mathbb{N}$, must decide between three modes of transportation: a known outside option ($o_t$), a new mode of transportation ($n_t$), and an integrated trip ($b_t$), which is a bundle consisting of the two types of transportation services. The amount consumed of each type of ride is then combined to produce total transportation services ($\tau_t$), according to the technology:
$$\tau_t = o_t + \gamma A^\phi b_t + A n_t \, , $$ where $A\geq 0$ is the relative quality of the new mode of transportation. The parameter $\phi \in (0,1)$ captures the fact that the new mode of transportation enters the composition of an integrated trip.\footnote{Indeed, the term $\gamma A^\phi b_t$ may be seen as a reduced form for the problem of choosing the composition of an integrated trip between the outside option and the innovation with Cobb-Douglas technology.}
Relative quality $A$ is unknown, but, at period $t$, the consumer has information $\mathcal{H}_t$ over it, where $\mathcal{H}_t$ is a sub-$\sigma$-algebra of $\mathcal{S}$, with $(S, \mathcal{S},\mathbb{P})$ being a probability space prescribing the uncertainty regarding $A$.\footnote{In other words, $A$ is a random variable defined on $(S, \mathcal{S},\mathbb{P})$.} We remain agnostic about the nature of uncertainty encoded in $\mathbb{P}$ -- it may reflect objective or subjective factors -- though we assume the consumer is an expected utility maximiser. Specifically, we assume that the consumer has preferences over both transporation services ($\tau_t$) and other forms of consumption ($c_t$), and that, given information at $t$, his expected utility at period $t$ is given by:
where $\lambda > 0$ and $\tau_t^*\geq 0$. In this setting, transportation enters utility as an essential good: the consumer has a target level of transportation ($\tau^*_t$), which may be thought as the amount required for receiving her income (or education), and any amount above or below that is undesirable to her.\footnote{The assumption of linearity in consumption may be seen to reflect that, in the effective range where the choice between transportation modes is made, the marginal utility of consumption of other goods is constant. See Argente2022 for a similar assumption of quasilinearity in outside consumption in a model for the demand for Uber rides under alternative payment methods.}
We are now ready to describe the dynamics of the choice problem. After accruing transportation services, the consumer receives a signal of the underlying quality $A$. We assume that this signal, $Z_t$, is given by $$Z_t = \psi(A) + \frac{1}{\bar{h}(b_t,n_t)} V_t, $$ where $\psi$ is a bimeasurable one-to-one map; $V_t|\mathcal{H}_t,A$ follows a stable distibution function $F$ with zero mean and unit variance:\footnote{In our setting, the distribution function $F$ is stable if, for every $c,c' \in \mathbb{R}$ and independent copies $Z, Z'$ distributed as $F$, there exists $c''$ such that $cZ+c'Z' \overset{d}{=} c'' Z$. } and $\bar{h}$ is a smooth increasing nonnegative function. This formulation captures the idea that, the more the user utilises the innovation, the more precise is the signal about its underlying quality. After observing the signal $Z_t$, information is then updated as:
$$\mathcal{H}_{t+1} = \sigma(\mathcal{H}_t, Z_{t}) \, .$$
Finally, given a discount factor $\beta \in (0,1)$, a consumer's value at $t$ is given by:
subject to the technology and budget constraints:
where $a_t$ is the agent's wealth, and $w_t$ is her (non-interest) income, and $i_t$ is the interest rate. The notation $\mathcal{H}_{t+1}(\bar{h}(b,n))$ represents the $\sigma$-algebra $\sigma(\mathcal{H}_{t}, \psi(A) + \frac{1}{\bar{h}(b,n)} V_t)$
The following proposition collects some properties of the learning model. In what follows, denote by $(a_t,c_t,b_t,n_t,o_t)_{t \in \mathbb{N}}$ the optimal choices of wealth and consumption:
The first part of the proposition shows that information is always valued in the model. In a finite horizon setting, a proof of this fact (for every period up to the terminal one) follows from extensions of Blackwell's comparison of experiments theorem Blackwell1953,DeOliveira2018 to general state spaces khan2020missing. In our infinite horizon setting, we alternatively rely on the stability assumption on the signal coupled with boundedness of income and the target to offer a direct proof.
The second part of the proposition relies on known results on martingales Durrett2019 to show that if, in an equilibrium, there are enough incentives to learning, inasmuch that the signal has always a sufficient degree of informativeness, then, asymptotically, the consumer is able to perfectly recover the underlying quality. An implication of this fact is that, in a world featuring consumers heterogeneous with respect to their beliefs and preferences, but a common underlying quality $A$ entering utilities, then, asymptotically, heterogeneity in consumption patterns due to different information vanishes.
The third part of the proposition is concerned with the effects of a one-period small change in the price of integrated rides on the demand for transportation. Item 3.(a) shows that, contemporaneously, this effect can be decomposed into a direct price effect, and indirect effects mediated by changes in the marginal value of learning. The latter may be interpreted as a change in the learning or experimentation motive for demand, as a one-period change in the price of integrated rides also alters incentives to accumulate information for the agent's future selves. The magnitude of each component is determined by the consumer's information regarding $A$ at the beginning of the period, as well as the technology parameters for bundled goods, $\gamma$ and $\phi$. Item 3.(b) shows that, in our model, extemporaneous effects on the demand for transportation are fully mediated by learning. Put another away, if a discount were not to induce any additional learning, then there would be no changes in the demand for transportation other than in the period where the discount takes place. Such phenomenon is driven by the essentiality nature of public transportation in the model.\footnote{Indeed, if there was complementarity between consumption and transportation in the model (e.g., the utility of consumption is given by $c_t^\alpha \tau_t^{1-\alpha}$ for some $\alpha \in (0,1)$), then intertemporal substitution effects in $c_t$ may induce changes in the demand for public transportation in other periods even if there was no change in the information acquired.}
We summarise below the two main predictions from the learning model, which will guide our empirical analysis:
We assess the predictions of our theory of learning in the context of transportation markets in S\ {a}o Paulo, Brazil. S\ {a}o Paulo is Brazil's largest municipality, with around 11.5 million inhabitants reported in the 2022 national Census census2022. Its transportation infrastructure includes a metro system extending short over 100km metro2022, nearly 200km of commuter rail lines connecting the municipality to its larger metropolitan area cptm2022, 1,300 bus routes sptrans2023 and 700 km of bus lanes sptrans2022, and 722 km of bicycle lanes cet2022. Since 2016, ride-hailing services are authorised to operate by the local government upon complying with regulations g12016,biderman2020regulating.
We rely on data from a randomised experiment which one of the authors conducted in the municipality of S\ {a}o Paulo Biderman2018 in partnership with a company that started operating a ridesharing service locally in 2016.\footnote{Before 2016, the company operated a local cab calling service. Cab rates are fixed by the local government, whereas ridesharing fees are dynamically set by the platform's algorithm.} In 2018, this service implemented an experiment among its registered users with an aim to understand the potential for its app rides being used to complete the first- or last-mile of a train/metro trip. The experimental design proceeded in two steps. In the first step, the company sent a message to a random sample of 60,000 of its users asking them to answer a survey in exchange for a 15 BRL\footnote{Around 4 USD at the 2018 exchange rate.} discount coupon upon completion. The initial survey collected socioeconomic data, as well as baseline information on commuting habits.\footnote{The variables collected in the baseline survey are presented in Table (ref) in the Online Appendix.} At this time, individuals were only aware of the 15 BRL coupon, with no further information on a future experiment being disclosed.
In a second step, a randomised experiment was conducted among the survey respondents. The experiment consisted in providing discounts to app rides starting or ending at a subway/train station during two weeks between the end of November and the beginning of December 2018. Respondents were randomly divided into three arms: (i) a control group (which was not informed about the experiment); (ii) a group eligible to two 20% discounts per day on rides starting/ending at a train/metro station,\footnote{The e-hailing company defined a virtual fence around the station. If the trip started or ended inside this fence, it was eligible to the discount.} limited to 10 BRL per ride; and (iii) a group eligible to two 50% discounts per day on rides starting/ending at a train/metro station, also limited to 10 BRL per ride. Announcement of the discounts proceeded as follows. On November 26th, 2018 (a Monday), users selected into treatment received SMS and push notifications informing them of the availability of two daily coupons for rides integrating with the train/metro system until November 30th (same week’s Friday). They also received daily reminders for the rest of the business week. On the next Monday (December 3rd), all users in these groups received similar notifications informing them that the discounts were also available for that whole week (until December 7th, Friday), again with subsequent daily reminders. As a consequence, users assigned treatments status were eligible for up to 200 BRL in discounts during these two weeks.
For each respondent, we observe their baseline survey answers, as well as the number of integrated (starting/ending at train/metro station) and non-integrated (total minus integrated) app rides on a fortnightly basis, from early 2018 until mid 2019. We view this setup as an ideal environment to assess the predictions of our theory. Indeed, it appears reasonable to model transportation as an essential good. Moreover, from the lenses of the model in (ref), we observe a one-off unanticipated discount to a bundled good (integrated rides), which upon utilization may provide information about a new mode of transportation (non-integrated rides). Finally, our time window, which spans over six months after the discount takes place, allows us to evaluate the long-run effects of this short-term discount.
To further improve the interpretation of our results, we restrict our analysis to respondents who reported either living, working or studying close to a train/metro station. We do so for two main reasons. First, it better allows us to interpret learning about $A$ as “learning about the relative quality of the innovation, vis-à-vis the train-metro option”. Indeed, users who do not report being usually close to a train-metro station may combine the subsidy with other modes of transportation (e.g. renting a bike at the station, or using another e-hailing service), which may induce learning about the relative quality of other forms of innovation, thus posing difficulties in the interpretation of results. Secondly, since, as discussed in the introduction, learning about the relative quality of ridesharing vis-à-vis public transportation may have nontrivial policy implications, estimates for this restricted subsample may be also more relevant from this perspective.
(ref) reports balance tests for the subsample of users close to a train/metro station. We note that groups are well-balanced. There is some evidence at the 10% level of there being differences in average daily expenses in transportation across groups, though these differences are small when compared to the reported average income.
The distribution of both integrated and non-integrated rides displays heavy-tails. This is evidenced by (ref), where we plot histograms for average pre-treatment biweekly rides. The average number of integrated biweekly rides is 0.08, whereas the largest reported value is 6.6. Similarly, the average number of non-integrated rides is 1.4, with the largest reported value being 20.7.
In settings with heavy tails, difference-in-means estimators of average treatment effects tend to perform poorly, with large sample sizes being required to achieve reasonable precision Lewis2015,Athey2023. Following Athey2023, we thus opt for semiparametric methods that, while constraining treatment effect heterogeneity, enable the construction of more efficient estimators.
Guided by our theoretical discussion, we propose a semiparametric model to approximate the effect of discounts on each type of ride. (ref) shows that the effect of a one-period unanticipated discount in integrated rides has both contemporaneous and extemporaneous effects. Contemporaneously, the magnitude of the impact depends on an agent's information set at the beginning of the period, as well as the technology parameters $\phi$ and $\gamma$. Extemporaneously, the effect is fully mediated by learning. Moreover, our results suggest that, in a world with heterogeneous agents, increased learning may not only change the average number of rides, but also its dispersion. Consequently, a model for treatment effects in our setting should: (a) allow for (some) heterogeneity with respect to prior information sets and the individual technology parameters for bundled goods ($\gamma$ and $\phi$); (b) allow for effects on both the mean and dispersion of outcomes.
In order to address (a), we consider separate models for respondents who reported being a regular train/metro user, and those that did not. We do so because these individuals may differ in their information regarding $A$ at the beginning of the experiment, as well as in their technology parameters $\gamma$ and $\phi$. We then proceed as follows. Let $b_t(d;u)$ denote the potential number of integrated rides at fortnight $t$ for an individual whose prior train/metro usage is $U = u \in \{0,1\}$, upon being assigned treatment $d \in \{0,20\%, 50\%\}$ on the fortnight starting on November 26th, 2018. For each $u \in \{0,1\}$ and $d \in \{20\%, 50\%\}$, we consider the following model:
Similarly, definining potential outcomes $n_t(d;u)$ for non-integrated rides, we consider the following model, for $u \in \{0,1\}$ and $d \in \{20\%, 50\%\}$.
Consistent with our discussion, models (ref) and (ref) allow discounts to change both the location and the dispersion of rides. Heterogeneity is unrestricted across prior train/metro usage status $u$ and time $t$, though within an usage-time cell, we restrict treated potential outcomes to follow a location-scale shift with respect to the non-treated potential outcome. Under such restrictions, efficient estimators can be constructed, as we discuss in (ref).
For the moment, suppose we have estimates $\hat{\alpha}_{b,t,d,u}$ and $\hat{\sigma}_{b, t,d,u}$ from model (ref). In this case, for $d \in \{20\%,50\%\}$, the population average effect in rides $\Delta(t,d,u) \coloneqq \mathbb{E}[b_t(d;u) - b_t(0;u)|U=u]$ may be estimated as:
where $p_{d,u}$ is the fraction of individuals with prior train/metro usage $u$ assigned to group $d$; and $\bar{b}_{t,d,u}$ is the average of integrated rides in fortnight $t$ among the subgroup with prior train/metro usage $u$ and assigned to arm $d$. Estimator (ref) relies on the semiparametric model (ref) to impute the relevant missing potential outcomes in the subgroup assigned to arms $d$ and $0$.\footnote{It is possible to construct an even more efficient estimator by combining models for different discounts and using these to input the potential outcome for treatment statuses $d$ and $0$ in the subgroup assigned status $d' \notin \{0,d\}$. While this may lead to a more precise estimator, we note that, in the presence of misspecification bias, it may lead to biases in one semiparametric model contaminating treatment effect estimates in another comparison, even if the model for the latter is correctly specified. This is why we opt for “self-contained” (involving a single-model) comparisons when estimating effects. See (ref) for further discussion.}
Our model also allows us to easily assess the effects of discounts on dispersion. Let $\Psi(t,d,u) \coloneqq \frac{\sqrt{\mathbb{V}[b_t(d;u)|U=u]}- \sqrt{\mathbb{V}[b_t(0;u)|U=u]}}{\sqrt{\mathbb{V}[b_t(0;u)|U=u]}}$ denote the relative change in dispersion induced by discount $d \in \{20\%,50\%\}$. We then estimate $\Psi(t,d,u)$ as:
In the next section, we discuss our approach to estimating $\hat{\alpha}_{b,t,d,u}$ and $\hat{\sigma}_{b, t,d,u}$.
Models (ref) and (ref) are particular parametrisations of the semiparametric models proposed by Athey2023. Consider a setting with a binary treatment and potential outcomes $Y(0)$ and $Y(1)$. Athey2023 analyse models of the form:
where $G$ is a known (up to $\theta_0$) function such that $y \mapsto G(y;\theta)$ is increasing, for every $\theta \in \Theta$. Given a random sample from the population and an experimental design that randomly assigns treatment to a subgroup of the sample, Athey2023 construct $\sqrt{N}$-consistent estimators for the parameter $\theta_0$ that asymptotically achieve the model's efficiency bound Newey1990,Bickel1993. Their estimator is constructed by adding a cross-fitted estimate of the model's efficient influence function to a first-step estimator. Estimation of the efficient influence function further requires (cross-fitted) estimates of the densities of $Y(0)$ and $Y(1)$.
In this paper, we propose an alternative estimator to the semiparametric model (ref) with convenient computational and statistical properties. Our estimation approach relies on L-moments, a robust alternative to standard moments. For a distribution function $F$ on the real line with quantile function $Q_F$, Hosking1990 defines the $r$-th L-moment as:
where $Q_F(u) \coloneqq \inf\{x \in \mathbb{R}:F(x)\geq u\}$, $u \in [0,1]$, is the $u$-quantile of $F$; and $P^*_{r}(u) = \sum_{k=0}^r (-1)^{r-k} \binom{r}{k} \binom{r+k}{k} u^k$ is a shifted Legendre polynomial. L-moments constitute an alternative to standard moments that is less sensitive to outliers. To see this, consider the second L-moment. In this case, Hosking1990 shows that $\lambda_2 = \frac{1}{2} \mathbb{E}[|Y_1-Y_2|]$, where $Y_1$ and $Y_2$ are independent random variables identically distributed to $F$. In contrast, we can show that the second centered moment is $\mathbb{V}[Y]= \frac{1}{2} \mathbb{E}[(Y_1-Y_2)^2]$, which is more sensitive to extreme values. L-moments are also known to characterise any distribution with finite first moment Hosking1990, a property not generally enjoyed by moments Billingsley2012.
Our proposed estimator is a semiparametric version of the generalised method of L-moments (GMLM) approach discussed in alvarez2023inference. Fix $R \in \mathbb{N}$, $R\geq p$, and let $\boldsymbol{P}_R(u) = (P_{0}^*(u), \ldots, P_{r-1}^*(u))'$. We propose to estimate $\theta_0$ as:
where $\hat{Q}_d$ is the empirical quantile function of the outcome in the group assigned treatment status $d \in \{0,1\}$; $0 \leq \underline{p} \leq \overline{p}\leq 1$ are possible trimming constants; and, for a vector $x \in \mathbb{R}^R$, $\lVert x \rVert_{W_R} = \sqrt{x'W_Rx}$, where $W_R$ is a $R\times R$ weighting matrix. Following the same logic of generalised method of moments estimators, the GMLM (ref) estimates $\theta_0$ by minimizing a weighted distance between the first $R$ empirical L-moments in the treatment group and the corresponding “model-implied” L-moments $\int_{\underline{p}}^{\overline{p}} G(Q_{0}(u); \theta) P^*_{r-1}(u) du$, where the unknown quantile function of the untreated potential outcome is replaced with its empirical counterpart. Our formulation explicitly accommodates for the possibility of trimming ($\underline{p}>0$ or $\underline{p}<1$), which may be employed in the presence of very extreme observations.
In the location-scale models (ref) and (ref), estimator (ref) can be easily computed. Indeed, in these cases, we have that, for any outcome $o \in \{b,n\}$, treatment $d \in \{20\%, 50\%\}$, prior train/metro usage $u \in \{0,1\}$ and fortnight $t$:
where
with $\hat{Q}_{o,t,s,u}$ being the empirical quantile function of outcome $o$ at period $t$, in the subgroup assigned treatment status $s$ and with prior train/metro usage $u$.
In Online (ref), we show that, in addition to a convenient computational formulation, our semiparametric GMLM estimator has attractive statistical properties. Online (ref) shows that, in an asymptotic framework where both the number of observations and the number of L-moments $R$ diverge, estimator (ref) is consistent and asymptotically normal. In this setting, we derive the optimal weighting matrix for L-moments and show how it can be conveniently estimated through a weighted bootstrap procedure. We also show that, in our setting, the minimum of the optimally-weighted objective function (ref), when multiplied by sample size, provides researchers with a specification test of overidentifying restrictions (J-test). In Online (ref), we show that, in the single-outcome and binary treatment setting of Athey2023, the optimally weighted L-moment estimator without trimming ($\underline{p}=0$ and $\overline{p}=1$) is asymptotically efficient.\footnote{We restrict our attention to the single outcome with binary treatment case because, when there are multiple outcomes or multiple treatments, it may be possible to construct a more efficient estimator by jointly estimating the parameters from different (sub)models. For example, in our empirical application, we note that equations (ref) and (ref) impose restrictions on the cross-correlation between integrated and non-integrated rides, contemporaneously as well as between different periods. Similarly, (ref) and (ref) jointly imply that for any individual in any experimental group, the two missing potential outcomes are identified. Therefore, by jointly estimating all the parameters for a given subgroup $U=u$, it may be possible to produce more efficient estimators.
We note, however, that a large drawback of this approach is that it would be much more sensitive to misspecification bias, since incorrect specification of the model (ref)/(ref) for a single tuple (outcome, period, discount, usage) would contaminate estimates for the remaining tuples with same usage. In order to avoid that, we thus separately estimate effects via (ref) for each tuple (outcome, period, discount, usage), and report p-values of the specification test in Online (ref) for the corresponding specification.
Finally, we note that, if the parametrisation given by (ref)/(ref) is considered individually for each tuple, then, by the argument in Online Appendix (ref), the estimator (ref) is efficient. Put another way, (ref) is asymptotically efficient for the semiparametric model that assumes its target submodel to be valid, but does not make any hypotheses about the submodels of the remaining tuples. } Inspired by the Synthetic Controls literature abadie2021using, Online Appendix (ref) proposes a method to select $R$ and, possibly, the trimming constants $(\underline{p}, \overline{p})$ by relying on pre-treatment data. Finally, Online (ref) shows that our proposed approach compares favourably to existing methods in a Monte Carlo exercise due to Athey2023.
We begin by reporting the effects of discounts on the average number of rides and their dispersion. We estimate the model parameters via (ref), with weights given by the estimator of the optimal weighting scheme discussed in Online (ref). The number of L-moments $R$ is selected according to the tuning procedure discussed in Online (ref). In light of the Monte Carlo exercise in Online (ref) showing limited gains to trimming, we do not allow for it, thus setting $\underline{p}=0<1=\overline{p}$. Parameter estimates are then used to compute the effects of discounts on average rides and dispersion, according to (ref) and (ref). Standard errors are computed via the delta-method.
Figures (ref) and (ref) plot estimates of the effects on the mean and dipersion, along with 95% confidence intervals, across the last two pre-treament fortnights (left of vertical dark line) and the post-treatment window, for each discount regime and outcome. We aggregate effects across prior train/metro usage group by averaging (ref) /(ref) across $u$, with weights given by the proportion of usage type $u$ in the sample. Consequently, estimates for average effects may be interpreted as estimating the population average treatment effect, whereas for the dispersion one may interpret them as the expected relative change in dispersion across usage groups.
Figure (ref) reports results for integrated rides. For the 50% discount, we observe a large and significant effect on average rides on the fortnight where the treatment is in place, and no significant effects thereafter. The magnitude of the contemporaneous effect is large, corresponding to 60% of the average in the control group at that fortnight. This result contrasts with the usually smaller estimates of own-price elasticities of demand for auto and public transit typically reported in the literature small2007economics. We also observe a large estimated effect for the 20% discount, corresponding to 30% of the control group average, though this effect is not statistically significant. For both discount regimes, there do not appear to be effects on the average nor on the dispersion of integrated rides after the discount is over.
Figure (ref) reports results for non-integrated rides. Our results show that, between two and three fortnights after the discount is over, we start observing reductions in both the average and the dispersion of integrated rides. These effects persist for over four months after the end of the discount, with reductions in dispersion between 10-20%. Effects are more pronounced in the 50% discount regime, though they are also detected in the 20% discount. By late March and early April, reductions in average rides in the 50% regime amount to 15% of the control group mean.
In order to shed more light on the underlying estimates, Tables (ref), (ref), (ref) and (ref) report estimated effects on the mean and dispersion of each outcome, for each discount and prior train/metro usage group. For comparison, we also report estimates of average effects based on difference-in-means estimators, which corrrespond to the efficient estimator in a nonparametric model that does not impose any structure on potential outcomes beyond finite second moments Newey1990. Nonparametric estimates of the effects on dispersion are based on the ratio of estimated standard deviations, minus one. For our semiparametric specifications, we also report the p-value of the overidentifying restrictions J-test discussed in Online (ref).
A few patterns are worth remarking. First, when estimating average effects, the semiparametric methods indeed lead to smaller standard errors than the difference-in-means estimators. Reductions can be as large as 45%. We also observe gains in precision when estimating effects on the dispersion of outcomes, with an average reduction in standard errors of 16%. Second, the location-scale parametrisation indeed appears to be adequate, as the overidentifying restrictions test only rejects the null at the 10% level for four specifications. In these four cases, inferences based on the nonparametric estimators lead to similar conclusions.\footnote{For completeness, Online Appendix (ref) replicates Figures (ref) and (ref), but replacing the semiparametric estimators used in each aggregation with the corresponding nonparametric estimators reported in Tables (ref), (ref), (ref), and (ref). Our main results remain unchanged.} Third, we note that the 50% discount induces contemporaneous increases in integrated rides for both the regular and non-regular train/metro user subpopulations. Moreover, for the 50% discount, we detect reductions in the mean and dispersion of non-integrated rides in both subpopulations, but especially among non-regular users. In several cases, however, standard errors are large. Finally, we note that, for the 20% discount, there is some evidence of reductions in the mean and dispersion of integrated rides among the nonregular user subpopulation, in the first fortnights after the discount is over.
Effects reported in the previous subsection are consistent with the following interpretation through the lenses of the model of learning in (ref). Upon increasing consumption of integrated rides, consumers learn that the relative quality of non-integrated rides was worse than they previously held (they update their beliefs to a more pessimistic value). This leads to a decrease in the average and dispersion of non-integrated rides, as consumers shift away from the innovation. However, since transportation is an essential good, consumers need to increase consumption of the remaining alternatives to satisfy their target. For bundled rides, the impact of relative quality $A$ is dampened by the technology parameter $\phi$; as a consequence, the negative effect introduced by belief updating is offset by the substitution motive, thus leading to the demand for integrated rides remaining approximately unchanged. Do also note that, as a consequence of these patterns, one would expect an increase in the demand for the (unobserved) outside option. The negative effects on non-integrated rides vanish after five months. We can interpret this pattern, within the model, through a catch-up in the beliefs in the control group; or, outside the model, through an improvement in the underlying quality $A$ or imperfect recall. Disentangling between these competing explanations would require direct measurements of subjects' assessments of relative quality, which we do not possess. We do note, however, that the trajectory of average nonintegrated rides in the control group is compatible with the catching-up mechanism, as we observe a steady decline in the posttreatment window (see (ref) in Online Appendix (ref)).
In order to better understand the workings of learning in our setting, we combine our experimental estimates with Proposition (ref) in order to quantify how much of the contemporaneous increase in integrated rides can be attributed to a learning or experimentation motive, i.e. to a desire to accumulate information for one's future selves. Recall that Proposition (ref) provides an additive decomposition of the contemporaneous effect of a one-period increase in bimodal rides into a direct price effect and a learning effect. Therefore, if we can reasonably calibrate the direct price effect, we may use our experimental estimates to recover the learning effect.
Proposition (ref) shows that the direct price effect depends crucially on the beliefs regarding $A$ at the beginning of the period, the technology parameters $\gamma$ and $\phi$, and the sensitivity-to-target parameter $\lambda$. To calibrate these quantities, we combine our experimental dataset with the 2017 Origin and Destination Survey (ODS). Funded by the São Paulo public metro company, the ODS collects a representative sample of the larger metropolitan area travel flows. In what follows, we describe in steps how we proceed with the calibration.
Calibration of prior beliefs regarding $A$: We assume that individuals hold a common prior regarding the relative quality $A$, which we calibrate using the 2017 ODS data on the quality of ridesharing trips. Specifically, we measure the quality of a trip in terms of travel distance in kilometres, and assume that the prior on $A$ is given by:
$$A = \frac{\xi \ell}{\bar{x} \bar{l}}, $$ where $(\xi,\ell)$ are independent log-normal random variables, with $\xi$ representing a trip's speed (travel distance over travel time), and $\ell$ representing travel time. The assumption decomposes travel quality onto two components: a measure of “productivity” $\xi$, and heterogeneity with respect to duration $\ell$. We calibrate the mean and variance of each component using the ODS data on ridesharing trips that do not integrate with a train/metro. We then normalise this quality by the average speed ($\bar{x}$) and duration ($\bar{l}$) of the outside option, i.e. nonridesharing trips.
Calibration of $\phi$ and $\gamma$: We assume that the quality of integrated rides, $\gamma A^\phi$, also admits a decomposition onto log-normals representing speed and duration. We calibrate these random variables from ODS data on ridesharing trips starting or ending at a train/metro. We then choose $\phi$ and $\gamma$ that match the first and second moments of this representation.
Calibration of $\lambda$: We allow $\lambda$ to be heterogeneous across individuals in our sample. In keeping with our interpretation of transportation services as an essential good, we assume that, if an individual chooses not to consume transport services, she loses her income. Considering a biweekly window for the model, we assume that:
$$\frac{\lambda_i}{2} \tau_i = \frac{w_i + \frac{w_i}{(1+r)}}{2} \, ,$$ i.e. the individual loses the present value of her monthly income $w_i$ if she does not meet her target. We assume the agent loses her income for two periods to capture the fact that transitions from and to employment are not frictionless Meghir2015. We observe $w_i$ for a subset of the users in the experimental sample. We then set the biweekly interest rates $r$ by considering an annual real interest rate of 4%. As for the target $\tau_i$, we impute it as follows. In the experimental sample, users were asked their average time spent going to work. We use this information to estimate, in the ODS, the user's expected travel distance. We multiply this value by 20, reflecting that in a fortnight an individual would come and go to work 20 times. We then normalise this quantiy by the productivity of the outside option. Having imputed $\tau_i$ and observing $w_i$ and $r$, we back out $\lambda_i$.
Table (ref) presents the results of our exercise. We report estimates and standard errors for both the sample average treatment effect on integrated rides,\footnote{The sample average treatment effect for regime $d \in \{20\%,50\%\}$ is defined as $\mathbf{T}_d = \frac{1}{|\mathcal{N}_d \cup \mathcal{N}_0|}\sum_{i \in \mathcal{N}_d \cup \mathcal{N}_0}(b_{\text{November, 26th}}(d;U_i) - b_{\text{November, 26th}}(d;U_i))$, where $\mathcal{N}_s$ is the set of individuals assigned to experimental group $s$ for which income and travel time data is available at the baseline. Subsample average effects are similarly defined. These parameters are estimated using the semiparametric model. Standard errors are computed using the Delta Method Athey2023.} and the percentage of this effect attributable to the learning motive. We note that, for the 50% discount, around $40\%$ to $50\%$ of the contemporaneous increase in integrated rides can be attributed to an experimentation motive. The estimated share is larger in the subsample of nonregular train/metro users -- precisely the subpopulation for which we detect the largest share of subsequent learning effects (see Tables (ref) and (ref)). As for the 20% discount regime, we do not detect significant effects neither on average rides nor on the share attributable to learning. In this case, estimates for the full sample are considerably more precise than those in subsamples, with the former indicating that the magnitude of the learning effect, if any, is likely to be small.
This paper showed that temporary subsidies to bundles may induce long-term changes in consumption behaviour due to learning about the quality of one of the constituent goods. We introduced a dynamic model where a consumer learns about an innovation on an essential good through repeated consumption. In the model, consumers leverage a one-period discount on a bundle that contains the innovation partly to increase the informational bequest left to her future selves, a mechanism we labeled the learning or experimentation motive for demand.
We then assessed the predictions of our theory by relying on data from a randomised field experiment conducted by a large ridesharing service in S\ {a}o Paulo. Given the heavy-tailed nature of demand in our setting, and guided by our theoretical discussion, we proposed a semiparametric specificication for treatment effects in our environment, and introduced an efficient estimator for our parametrisation by considering a “plug-in” generalised method of L-moments estimator.
Our semiparametric approach enables us to uncover a range of patterns that are broadly consistent with our theory of learning about an essential good. A ten-day discount in app rides integrating with a train/metro station generates persistent negative effects in the mean and dispersion of the demand for non-integrated rides. These effects last for several months, which, given the essentiality nature of transportation services, is interpreted as further evidence of learning. A simple calibration exercise then shows that, for the largest discount regime, around 40% to 50% of the contemporaneous increase in integrated rides may be attributed to the experimentation motive.
Insofar as the outside option in our analysis may be interpreted as public transport (PT), our results would indicate that temporary subsidies to modal integration may generate persistent increases in the demand for PT. Indeed, since the average demand for integrated rides does not decrease with the discount, and that, given the estimates from our analysis, our model predicts an increase in the demand for the outside option after the discount is over; our results would reveal that, by improving the relative assessment on the quality of PT, temporary subsidies on modal integration may lead to persistent increases in PT takeup. From a policy perspective, this suggests that such incentives may be a useful tool in partially offsetting secular declines in PT usage.
Even if the effects of learning are not taken into account, a simple calculation based on the point estimates for the 50% discount and the decomposition between learning and direct effects reported in Table (ref) suggests that the contemporaneous, direct (not driven by learning) own-price-elasticity of the demand for integrated rides to be at least 0.6.\footnote{Using the information in Table (ref), we calculate elasticities as: $$-\frac{(1-\text{share\_due\_to\_learning})\text{sample\_avg\_effect}}{0.5\text{mean\_control}}\, .$$ This calculation hinges on the assumption that all subsidised rides received a 50% discount. Given that there was a cap of 10BRL in the discounts (which we do not observe, as price data is not available to us), it may thus be seen as a lower bound to the actual elasticity.} This is much larger than values between 0.3 and 0.4, the typical range used as a rule-of-thumb for own-price-elasticities of the demand for PT small2007economics. Public policies aiming at increasing the share of public transit in total trips typically rely on subsidising tariffs. Our calculations suggest that an alternative policy, to integrate e-hailing with public transit through a discount, may induce larger public transport takeup. The increase in takeup could help paying for the large fixed cost in operating PT systems. Last-mile e-hailing rides would complement the system in a potentially more efficient manner, as low-density routes may be better served through ridesharing than through more costly (and polluting) mass transit alternatives.
There are several venues of research which stem from this work. From a policy perspective, it would be interesting to design experiments that directly measure the outside options, so as to properly quantify the PT takeup induced by both the direct effect and learning. From a market design perspective, it would be important to understand what is the optimal pricing policy of a firm when learning-from-bundling is taken into account. Finally, from the econometric viewpoint, it would be relevant to understand more generally the conditions for semiparametric efficiency of “plug-in” GMLM estimators, for example by extending arguments from the generalised method of moments literature Ackeberg2014.
\enddoc@text\let\enddoc@text\relax
\setcounter{page}{1}