Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,559 characters · 25 sections · 44 citation commands
Personalized Pricing with Invalid Instrumental Variables: Identification, Estimation, and Policy Learning
\if00 {
} \fi
In the era of Big Data and artificial intelligence, business models and decisions have been changed profoundly. The massive amount of customer and/or product information offers an exciting opportunity to study personalized pricing strategies. Specifically, based on the information collected from past selling seasons, sellers can leverage powerful machine learning tools to discover their customers' preferences and offer an attractive personalized price for each customer to maximize their revenue.
This problem can be formulated as an offline policy learning problem for continuous treatment space. In particular, offline data often consist of customer/product information, the offered price, and the resulting revenue. Our goal of personalized pricing is to leverage such data to discover an optimal data-driven pricing strategy for each $X$ that maximizes the overall revenue.
Because we have no control over the collection of the offline data, one major challenge of this task is that there may exist unmeasured confounders besides the offline data, which could possibly result in endogeneity. Endogeneity typically hinders us from identifying an optimal pricing decision using offline data. Thus, using standard policy learning methods may lead to suboptimal pricing decisions.
In the literature on causal inference and econometrics, instrumental variable (IV) models are commonly used to account for unmeasured confounding in identifying the causal effect of treatment. This task is closely related to personalized pricing because evaluating a particular pricing strategy is almost equivalent to its causal effect estimation. Therefore, an IV model is a promising solution for addressing the endogeneity in finding an optimal pricing strategy. A valid IV is a pretreatment variable independent of all unobserved covariates, and only affects the outcome through the treatment. Meanwhile, it requires that the variability of IV can account for that of the unobserved covariates. Prominent examples include using the season of birth as an instrument for understanding the impact of compulsory schooling on earnings in education angrist1991does, estimating the spatial separation of racial and ethnic groups on the economic performance using political factors, topographical features, and residence before adulthood as instruments in social economics cutler1997ghettos, estimating the effect of childbearing on labor supply using the parental preferences for a mixed sibling-sex composition as an instrument in labor economics, justifying the Engel curve relationship of individual household's expenditure on the commodity demand using individual's expenses on nondurables and services as an instrument in microeconomics blundell2007semi, and leveraging genetics variants as instruments for investigating the causal relationship between low-density lipoprotein cholesterol on coronary artery disease in medical study burgess2013mendelian.
We study offline personalized pricing under endogeneity via an IV approach. Instead of assuming a valid IV, we use an invalid IV that can potentially have an additional direct effect on the outcome. Relying on the structural models of revenue and price, we establish the identifiability condition of an optimal pricing strategy given observed covariates with the help of invalid IVs. Based on the identification, which can be formulated as a problem of solving conditional moment restrictions with generalized residual functions, we develop an adversarial min-max estimator and learn an optimal pricing decision from the offline data. We call our proposed policy learning method PRINT: Personalized pRicing using Invalid iNsTruments. Most existing literature focuses on developing methods for either a discrete treatment space with some valid/invalid IVs or a continuous treatment space with a valid IV. While lewbel2012using,tchetgen2021genius also considered the causal effect estimation for a continuous treatment using an invalid IV, their IV models do not allow for the causal heterogeneity, e.g., the interaction effect of the price with covariates. This cannot serve our purpose of finding an optimal personalized pricing strategy.
One of our motivating examples is the pricing problem of an auto loan company. (Later, we will conduct extensive numerical experiments on this problem using a real dataset.) Customer lending is a prominent industry in which personalized prices (i.e., lending rates) are both socially acceptable and in current practice, albeit at varying degrees of granularity ban2021personalized. The norm of price negotiation, high variation in customer willingness to pay, low cost of bargaining, and other considerations provide tremendous opportunities for offering personalized lending rates for customers in order to maximize the profit phillips2015effectiveness.
In this example, the offline data consist of customer information (e.g., FICO score, loan amount, loan term, living state), the offered loan price calculated by net present value, and the resulting revenue determined by the final contract result (accepted or not). However, when recording the information of past deals in the offline data, some local information, such as the operating costs of the lender, the competitor's rate on individual deals, and some unknown customer demographics, may not be available, which hinders the decision maker from designing an optimal pricing strategy based on the available covariates information. Thus, it is essential to devise an offline learning algorithm for personalized pricing under endogeneity, which learns the pricing strategy based on available covariates (i.e., a mapping from covariates to prices).
To account for unmeasured confounding, the loan rate, or the so-called APR (annual percentage rate), can be served as an instrument variable $G$ for dealing with the endogeneity of the price of the loan, because it is strongly relevant to the price and satisfies specific properties. We also note that using an APR as an IV has been adopted in the literature (e.g., blundell1992credit). However, a caveat is that the loan rate may inevitably affect the demand directly and, subsequently, the revenue of the loan company, which breaks the exclusion restriction to be a valid IV. In fact, in many problems, the IV exclusion restriction is hard to be verified. To overcome this difficulty, it is necessary for us to develop a new policy learning approach using an invalid IV.
We study the problem of offline personalized pricing under endogeneity. Our contribution can be summarized four-fold. First, we develop a novel policy learning method for continuous treatment space using invalid IVs. Our identification using invalid IVs for an optimal pricing strategy relies on two practical non-parametric models of revenue and price. To the best of our knowledge, this is the first work studying policy learning for continuous treatment space under unmeasured confounding. Second, we generalize the causal inference literature on treatment effect estimation under unmeasured confounding. In particular, the existing literature is predominantly focused on discrete treatment settings, where various approaches using a valid IV are developed. Much less attention has been paid to dealing with continuous treatment. While there is also a stream of recent literature studying causal effects under invalid IVs, the significant works still concentrate on the discrete treatment setting. Our identification using an invalid IV for the effect of a continuous treatment fills the gap of causal inference literature, which could be of independent interest. The key step for establishing the identification is to impose orthogonality conditions in terms of the high-order moments between the effect of all covariates related to the price on the outcome and the effect of that on the price, while the degree of unmeasured confounding is not restricted. Third, we establish an asymptotic regret guarantee for our policy learning algorithm in finding an optimal pricing decision, based on a newly developed adversarial min-max estimator for solving conditional moment restrictions with generalized (non-linear) residual functions. This adversarial min-max estimator is motivated by solving a zero-sum game that can incorporate flexible machine learning models. Lastly, compared with two baseline methods, we demonstrate our method's superior performance via extensive simulation studies and a real data application from an auto loan company.
Since the seminal work of manski2004statistical, there has been a surging interest in studying offline policy learning in economics, statistics, and computer science communities such as qian2011performance,dudik2011doubly,zhao2012estimating,chen2016personalized,kitagawa2018should,cai2021jump,biggs2022convex,qi2022offline and many others. However, most existing works rely on the unconfoundedness assumption, which is hardly satisfied in practice. To remove the effect of possible endogeneity, practitioners often collect and adjust for as many covariates as possible. While this may be the best approach, it is often very costly and even infeasible as we have no control over offline data collection. To address this limitation, more recently, various policy learning methods under unmeasured confounding have been proposed, such as using a binary and valid IV for a point or partially identifying the optimal policy cui2021semiparametric,qiu2021optimal,han2019optimal,pu2020estimating,stensrud2022optimal, using a sensitivity model for policy improvement kallus2020confounding, and leveraging the proximal causal inference qi2022proximal,miao2022off,wang2022blessing,shen2022optimal. However, none of the aforementioned works studies the policy learning for continuous treatment space under endogeneity.
Our work is also closely related to IV models, which have been extensively studied in the literature on causal inference and econometrics. See angrist1995identification,ai2003efficient,newey2003instrumental,hall2005nonparametric,chen2011rate,chen2018optimal,darolles2011nonparametric,blundell1992credit,wang2018bounded for earlier references. A typical assumption in the aforementioned literature is the existence of a valid IV that satisfies (i) independence from all unobserved covariates $U$, (ii) the exclusion restriction that prohibits the direct effect of IVs on the outcome, and (iii) correlated with the endogenous variable. Tremendous efforts have been made to develop statistical and econometric methods to account for the possible violation of these assumptions. See staiger1994instrumental,stock2000gmm,stock2002survey,chao2005consistent,newey2009generalized and many others for relaxing (iii), and lewbel2012using,kang2016instrumental,guo2018confidence,tchetgen2021genius,sun2021semiparametric for relaxing (i) and (ii). In particular, kang2016instrumental,guo2018confidence,kolesar2015identification,windmeijer2019use considered the multiple IV setting and restricted to some specific parametric models. In contrast, lewbel2012using,tchetgen2021genius, which are closely related to our proposal, mainly focused on the discrete treatment setting and only considered the constant causal effect for continuous treatment setting. Therefore, none of the existing works studies the causal identification with heterogeneity under continuous treatment using an invalid IV, which is a distinct aspect of our paper.
In this section, we introduce the problem of personalized pricing under the framework of causal inference. We also illustrate the challenges of finding an optimal pricing strategy in the observational study due to endogeneity.
Let $P$ be the price of a product that takes values in a known and continuous action space $\mathcal {P} = [p_1, p_2]$ with $0 \leq p_1 \leq p_2$. Define the potential revenue under the intervention of $P=p$ as $Y(p)$ for $p \in \mathcal {P}$. In the counterfactual world, $Y(p)$ is a random variable of the revenue had the company used the price $p$ for their product. Denote $X$ as the observed $q$-dimensional covariate associated with the product that belongs to a covariate space $\mathcal {X} \subseteq \mathbb{R}^q$. A personalized pricing policy $\pi$ is determined by the covariate $X$, which is a measurable function mapping from the covariate space $\mathcal {X}$ into the action space $\mathcal {P}$. Then the potential revenue under a policy $\pi$ is defined as $Y(\pi)$. For any pricing strategy $\pi$, we use the expected revenue (also called policy value) to evaluate its performance, i.e.,
Finally, the goal of personalized pricing is to find an optimal policy $\pi^\ast$, such that
where $\Pi$ is a class of policies depending on $X$. However, since for each instance of the random tuple $(X, P, Y)$, we only observe one $Y$ corresponding to the price $P$, but not other $Y$ and $P$, the joint distribution of $(X, P, \{Y(p)\}_{p \in \mathcal {P}})$ is impossible to learn without any assumptions. Therefore, identification conditions are needed for learning $\mathcal {V}(\pi)$ from the observed data. We first consider the following three standard causal assumptions.
Assumption (ref)(ref) ensures that the observed $Y$ matches the potential revenue under the intervention $P$. Assumption (ref)(ref) guarantees that each pricing decision has a chance of being observed. The unconfoundedness assumption, i.e., Assumption (ref)(ref), indicates that by conditioning on $X$, there are no other factors that confound the effect of the price $P$ on the revenue $Y$. Under Assumptions (ref)(ref) and (ref)(ref), one can show that for each $\pi\in\Pi$, $$ \mathcal {V}(\pi) = \mathbb{E}[Q(X, \pi(X)], $$ where $Q(x, p) = \mathbb{E}[Y \, | \, X = x, P = p]$, and the expectation is taken over $X$. Together with Assumption (ref)(ref), $\mathcal {V}(\pi)$ can be nonparametrically identified by the observed data. Then $\pi^\ast$ satisfies $\pi^\ast \in \operatorname*{arg\,max}_{\pi \in \Pi} \mathbb{E}[Q(X, \pi(X)]$, whose explicit form is $\pi^\ast(X) \in \operatorname*{arg\,max}_{p \in \mathcal {P}} Q(X, p)$, almost surely.
In practice, the unconfoundedness assumption (i.e., Assumption (ref)(ref)) cannot be ensured without further restriction on the data-generating procedure such as an ideal randomized experiment. The failing of Assumption (ref)(ref) incurs non-identifiability issue and thus could lead to a seriously biased estimation for the policy value $\mathcal {V}(\pi)$. Consider the following toy example of a revenue model for illustration that
where $U$ is some unmeasured covariate with $\mathbb{E}[U \, | \, X] = X^\top \gamma$ for some unknown parameter $\gamma$; $\beta$ is the parameter of interest, and $\varepsilon$ is some random noise such that $\mathbb{E}[\varepsilon \, | \, X, U, P]=0$ almost surely. Suppose that we aim to evaluate a policy $\pi_0(X) \equiv 1$, i.e., always assigning the price $P =1$, then one can show that
Due to the unobserved factor $U$, we cannot identify the parameter of interest $\beta$, which is the effect associated with $\pi_0$, based on the observed data. Meanwhile, if one carelessly implements the previous approach by assuming unconfoundedness, then, since $\mathbb{E}\left[U \, | \, X, P = 1\right] \neq \mathbb{E}\left[U \, | \, X\right]$ in general, we have
In the causal inference, to account for unmeasured confounding, IV models are widely used in identifying the average treatment effect angrist1996identification. It is often assumed that there exists an IV, denoted by $G$, such that
for some function $\Upsilon$ and random noise $\widetilde \varepsilon$. A valid IV satisfies the following three conditions:
It is well-known that Assumption (ref) is sufficient for a valid statistical test of no individual causal effect of $P$ on $Y$, but not to point identify the average treatment effect, e.g., $\mathbb{E}[X^\top \beta]$ in (ref). An additional assumption is often needed to achieve the latter goal. However, even identifying the average treatment effect does not suffice in personalized pricing because to learn an optimal pricing strategy $\pi^\ast$, we need to identify the causal heterogeneity effect. For example, in (ref), one can find $\pi^\ast$ via solving
and the key is then to estimate the function $X^\top \beta$. Therefore, compared with the standard causal inference using a valid IV, this posits an additional challenge.
Furthermore, when $P$ is continuous, existing literature often assumes an additively separable structural model such as (ref) and further restricts that $G \protect\mathpalette{\protect\independenT}{\perp} \varepsilon$. The identification of $X^\top \beta$ in the toy example is then given by solving a conditional moment restriction $\mathbb{E}[Y + P^2 - X^\top \beta \times P \, | \, X, P, G] =0$ ai2003efficient,newey2003instrumental. Later on, researchers found that the separable structural equation could be dropped by restricting the dimensionality and heterogeneity of $U$ in affecting $Y$. See, for example, chernozhukov2007instrumental. While significant efforts have been made recently to further relax the condition on the outcome/revenue model in terms of $U$, none of them consider the circumstance where the exclusion restriction in Assumption (ref) or $G \protect\mathpalette{\protect\independenT}{\perp} \varepsilon$ fails. For example, in our auto loan application, the APR of a loan is used as the IV, which may have an unavoidable direct effect on the eventual revenue.
Given these challenges, in the following section, we consider invalid IVs and develop a novel identification for an optimal pricing strategy from the observational data.
In this section, we present an identification result using invalid IVs, denoted by $G$, which can be multi-dimensional, for finding $\pi^\ast$ defined in (ref) under endogeneity. Our result is based on realistic non-parametric models for the price and revenue, together with the restriction on the directions and strength of instruments in affecting the price and revenue.
In the following, we introduce our non-parametric revenue and price models in the presence of unmeasured confounders $U$, which could be multi-dimensional as well. Denote the space of $U$ as $\mathcal {U}$. We assume the following structural equation models for our data-generating process of a random tuple $(X, U, G, P, Y)$ that
For simplicity, we assume that $\beta_{p, 2}(U, X) \leq -c$ almost surely for some constant $c >0$. We term (ref) and (ref) as revenue and price models, respectively. In the revenue model, we consider a quadratic model of the price $P$ on the revenue $Y$, which is practical when considering the linear demand function bastani2022meta. The unknown coefficient functions $ \beta_{p, 1}(U, X)$ and $ \beta_{p, 2}(U, X)$ represent the linear and quadratic effects of the price $P$ on the expected revenue $Y$. In addition, the coefficient functions $\beta_g(U, X, G)$ and $\beta_{u, x}(U, X)$ denote the generic effect of $(X, U, G)$ on the revenue $Y$. Specifically, $\beta_g(U, X, G)$ characterizes the interaction effect of $G$ and $(X, U)$. We remark that without any additional assumptions, both $\beta_g(U, X, G)$ and $\beta_{u, x}(U, X)$ cannot be identified due to the unobserved $U$. However, the non-identifiability of these two functions does not necessarily hinder from finding $\pi^\ast$ as they are irrelevant to the price $P$ in (ref). For ease of presentation, the revenue model (ref) considered here rules out the interaction effect between the instrumental variable $G$ and the price $P$ on the revenue $Y$. Our method can be naturally extended to that scenario as well. In comparison, clearly, our revenue model (ref) is more general than the parametric and additive models used in the standard instrumental variable regression. In addition, we allow for the direct effect of $G$ on $Y$, which is typically not allowed in most existing literature on causal inference and econometrics.
Our price model (ref) is very flexible and describes a distributional aspect of the prices in our offline data, stemming from all relevant variables $(X, U, G)$. Note that it is unnecessary to identify nuisance functions $\alpha_g(U, X, G)$ and $\alpha_{u, x}(U, X)$ because they are irrelevant to the price $P$ and finding the optimal pricing strategy.
Due to the unmeasured confounding $U$, in our offline data, for each instance, we can only observe a sample of a random tuple $(X, G, P, Y)$. Under this model setup, our goal is to find the optimal personalized pricing strategy that maximizes the expected revenue. We do not consider $\pi^\ast$ that depends on $G$, as indicated by model (ref), there is no interaction effect between $G$ and $P$ on the expected revenue $Y$. Then we have the following proposition.
Let $\beta_{p, 1}(X) = \mathbb{E}\left[\beta_{p, 1}(U, X) \, | \, X\right]$ and $\beta_{p, 2}(X) = \mathbb{E}\left[\beta_{p, 2}(U, X) \, | \, X\right].$ Since $U$ is not observed in our data, without any assumptions, $\beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$ cannot be uniquely identified by the observed data non-parametrically. The non-identification issue indicates that there may exist two different expected revenues under the distribution of the observed random tuple $(X, G, P, Y)$. Directly applying supervised learning from $Y$ on $(X, G, P, P^2)$ will lead to biased estimation of $\mathbb{E}\left[\beta_{p, 1}(U, X) \, | \, X\right]$ and $\mathbb{E}\left[\beta_{p, 2}(U, X) \, | \, X\right]$, and thus the resulting estimated policy will be sub-optimal. See the toy example (ref) in the previous section for illustration.
We impose the following identification assumptions on $\beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$ by leveraging an invalid IV $G$.
Assumption (ref)(ref) ensures that the IV $G$ is correlated with the price $P$ given the observed covariates $X$, which is mild. This is a typical assumption for IVs approach so that we can use for adjusting the unmeasured confounding. Assumption (ref)(ref) essentially requires that there is no unmeasured confounding to infer the effect of $G$ on $Y$ by adjusting the observed confounders $X$. This holds for example, when $U$ is some private information owned by the competitor in our auto loan example. Assumption (ref)(ref) is a technical condition used to ensure that there are no common effect modifiers resulting from the unobserved covariates $U$ in both Models (ref) and (ref). Intrinsically, orthogonality conditions (ref) -- (ref) impose further strength requirements of the IV $G$ and covariates $X$ such that the effects of unmeasured confounding $U$ on the revenue are orthogonal to the conditional pricing moments $\mathbb{E}[\texttt{poly}_3 (P)\, | \, G,X,U]$ with $\texttt{poly}_3 (P)$ being any polynomial of price $P$ up to the third order. Note that we do not impose any restriction on the relationship between $\beta_{u, x}(U, X)$ and $\alpha_{u, x}(U, X)$, and hence the effect of unmeasured confounding can be arbitrarily large.
Below we provide a sufficient condition so that Assumption (ref)(ref) holds.
Assumption (ref) basically states that there exist two independent variables $U_1$ and $U_2$, which have separate effects on the price and revenue, respectively. In the context of the car loan example, $U_1$ could be the competitor’s rate on individual deals, whereas $U_2$ could be the operating costs of the lender. Then we have the following proposition.
Now, by the aforementioned assumption, we establish our identification results for $ \beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$ using our offline data. This relies on the following two lemmas.
We then have the following theorem for identifying $\beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$.
Given the identification result, we are able to perform policy learning. This basically consists of two steps. The first step is to use the offline data to estimate $\beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$ based on Theorem (ref), after which we can obtain the optimal policy $\pi^\ast$ via Equation (ref).
In this section, we discuss how to leverage the offline data to estimate $\beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$, and perform policy optimization. To estimate $\beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$, one can first estimate nuisance parameters $\Omega_1$-$\Omega_3$ and $\Upsilon_1$-$\Upsilon_3$, and then construct estimators based on Theorem (ref). However, one cannot directly implement supervised learning techniques for obtaining these nuisance parameters as they involve a nested conditional expectation structure. For example, the response variable in $\Omega_1$ is not directly observed, and one has to estimate $\mathbb{E}[G\, | \, X]$ and $\mathbb{E}[P \, | \, G, X]$ first, which will induce additional errors in estimating $\Omega_1$. In the following, we formulate the estimation problem as solving conditional moment restrictions with generalized residual functions, and develop an adversarial min-max approach to simultaneously estimating $\beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$.
Denote relevant nuisance parameters as $h = (h_1, \cdots, h_6)^\top$, where
and let
and let $\wt W_2 = (w_7, w_8)^\top$, where
In particular,
with $\text{Vec}(\rho) = (\rho_1, \cdots, \rho_6)^\top$.
Then we have the following lemma that characterizes the property of all nuisance parameters. To lighten the notation, let $Z = (Y,X,G,P)$, $\alpha_0 = (\beta_{p, 1}, \beta_{p, 2}, h_1, \dots, h_6)$ and $W(Z; \alpha) = \left(\text{Vec}^\top(\wt W_1), \text{Vec}^\top(\wt W_2) \right)^\top$ for any generic $\alpha$.
Equation (ref) is called a non-parametric non-linear instrumental variable problem chen2012estimation, where $Z = (Y, X, P, G)$ are endogenous variables, and $X$ is an instrumental variable. The IV $X$ in (ref) is for estimating all the nuisance parameters $\alpha$, and the IV $G$ in the revenue model (ref) is for identifying $\beta_{p, 1}(X)$ and $\beta_{p, 2}(X)$ in the presence of the unobserved confounder $U$. While both of them are called IVs, they serve different purposes. The non-parametric nonlinear IV problem is much more challenging than the standard non-parametric IV regression, i.e., $W$ is a linear function of $\alpha$, which has been extensively studied in statistics and econometric literature. In contrast, the general non-parametric nonlinear IV problem is much less studied theoretically, where chen2012estimation,chen2015sieve only studied this problem under the linear sieve model. Given a wide range of machine learning approaches, it is essential to study this estimation problem under flexible non-parametric function classes such as reproducing kernel Hilbert spaces (RKHSs), neural networks, high-dimensional linear models, etc. Motivated by these and also dikkala2020minimax,bennett2020variational, we reformulate Equation (ref) into an unconditional moment restriction via a min-max criterion, based on which we estimate $\alpha_0$ via solving a zero-sum game.
Specifically, suppose that we are given $n$ independent and identically distributed samples $\mathcal {D}_n = \left\{Z_i = (X_i, G_i, P_i, Y_i) \right\}_{i = 1}^n$, and a initial guess $\widetilde{\alpha}$ of $\alpha_0$. We propose to obtain $\widehat{\alpha}_n$ as an estimator of $\alpha_0$ via solving
where $ \Psi_n(\alpha,f) = \frac{1}{n} \sum_{i=1}^n W(Z_i, \alpha)^{\top}f(X_i), $
and $\mathcal {H}$ and $\mathcal {F}_i$ are some user defined function spaces. Examples of $\mathcal {F}_i$ and $\mathcal {H}$ include RKHSs, random forests, neural networks, and many others. The norm associated with $\mathcal {F}$ is defined as $\|f\|_{\mathcal {F}}^2 = \sum_{i = 1}^{d_W} \|f_i\|_{\mathcal {F}_i}^2$, where $\|\bullet\|_{\mathcal {F}_i}$ is some functional norm, and $\|\bullet\|_{\mathcal {H}}$ is the norm associated with the space $\mathcal {H}$, and $\lambda_n,\mu_n>0$ are regularization parameters. In addition, $\|f\|_{\widetilde{\alpha},n}^2 = \frac{1}{n} \sum_{i=1}^n \{f(X_i)^{\top}W(Z_i;\widetilde{\alpha})\}^2$, where $W(Z_i;\widetilde{\alpha})$ is used to balance the weights of conditional moment restrictions. The validity of using this min-max criterion in finding $\alpha_0$ can be verified by the following lemma.
After obtaining $\widehat \alpha_n$, we compute an estimator of $\pi^\ast$ by policy learning that
Notice that $\widetilde{\alpha}$ may not serve as an accurate initial guess of $\alpha_0$, which may lead to an unsatisfactory estimator $\widehat{\alpha}_n$. However, we can update it iteratively (see Algorithm (ref)).
In this section, we evaluate the performance of the learned policy $\widehat{\pi}$ by (ref). We show that with a given function space ${\mathcal{H}}$ that contains $\alpha_0$ and a proper adversary space ${\mathcal{F}}$ containing a function that can well approximate the projected generalized residual functions $m(X;\alpha)\triangleq\mathbb{E}[W(Z;\alpha)\, | \, X]$, the min-max estimator $\widehat{\alpha}_n$ obtained by (ref) is consistent to $\alpha_0$ in terms of ${\mathcal{L}}^2$-error. Based on the consistency of $\widehat{\alpha}_n$, we further obtain an asymptotic rate of $\widehat{\alpha}_n$ measured by a pseudometric $\|\bullet\|_{ps,\alpha_0}$ defined in (ref) below, in terms of the critical radii of spaces ${\mathcal{H}}$ and ${\mathcal{F}}$. Finally, we show that an asymptotic bound for regret of the learned policy $\widehat{\pi}$ by imposing a link condition for the pseudometric and ${\mathcal{L}}^2$ norm.
We start with some preliminary notations and definitions. For a given normed functional space ${\mathcal{G}}$ with norm $\|\bullet\|_{{\mathcal{G}}}$, let ${\mathcal{G}}_M = \left\{ g\in{\mathcal{G}}: \|g\|_{{\mathcal{G}}}^2 \leq M \right\}$. For a vector valued functional class ${\mathcal{G}} = \left\{ {\mathcal{X}} \rightarrow {\mathbb{R}}^d \right\}$, let ${\mathcal{G}}|_k$ be the $k$-th coordinate projection of ${\mathcal{G}}$. Define the population norm as $\|g\|_{p,q} = \|g\|_{{\mathcal{L}}^p(\ell^q,P_X)} = [\mathbb{E}\|g(X)\|_{\ell^q}^p]^{1/p}$ and the empirical norm as $\|g\|_{n,p,q} = \|g\|_{{\mathcal{L}}^p(\ell^q, X_{1:n})} = [\frac{1}{n}\sum_{i=1}^n \|g(X_i)\|_{\ell^q}^p]^{1/p}$.
Our main results rely on some quantities from empirical process theory wainwright2019high. Let ${\mathcal{F}}$ be a class of uniformly bounded real-valued functions defined on a random vector $X$. The local Rademacher complexity of the function class ${\mathcal{F}}$ is defined as
where $\left\{ \epsilon_i \right\}_{i=1}^n$ are i.i.d. Rademacher random variables taking values in $\left\{ -1,1 \right\}$ with equiprobability, and $\left\{ X_i \right\}_{i=1}^n$ are i.i.d. samples of $X$. The critical radius $\delta_n$ of ${\mathcal{G}} = \left\{ g:{\mathcal{X}} \rightarrow {\mathbb{R}}^d, \|g\|_{\infty,2}\leq 1 \right\}$ is the largest possible $\delta$ such that $ \max_{k=1,\dots,d}{\mathcal{R}}_n(\delta, \text{star}({\mathcal{G}}|_k)) \leq \delta^2, $ where the star convex hull of class ${\mathcal{F}}$ is defined as $\text{star}({\mathcal{F}}) \triangleq \left\{ rf: f\in{\mathcal{F}}, r\in[0,1] \right\}$.
Without loss of generality, suppose that there exist constants $A,B$ such that $\|\alpha\|_{\infty,2}\leq 1$ and $W(\bullet,\alpha)\in[-1,1]^{d_W}$ for all $\alpha\in{\mathcal{H}}_A$, and $\|f\|_{\infty,2}\leq 1$ for all $f\in{\mathcal{F}}_B$. Furthermore, for any $\alpha\in{\mathcal{H}}$, let $\Psi(\alpha,f) \triangleq \mathbb{E} [W(Z,\alpha)^{\top} f(X)]$ and $\|f\|_{\alpha}^2 \triangleq \mathbb{E} [ \{f(X)^{\top} W(Z;\alpha)\}^2 ]$, which are populational analogues of $\Psi_n(\alpha,f)$ and $\|f\|_{n,\alpha}$, respectively. By letting $\Sigma_{\alpha}(X) \triangleq \mathbb{E} [ W(Z;\alpha) W(Z;\alpha)^{\top} \, | \, X ]$, we have that $\|f\|_{\alpha}^2 = \mathbb{E} \left[ f(X)^{\top}\Sigma_{\alpha}(X) f(X) \right]$. In addition, let $\Phi(\alpha) = \mathbb E[m(X;\alpha)^{\top} \{\Sigma_{\alpha_0}(X)\}^{-1}m(X;\alpha)]$.
In this section, we introduces some preliminary assumptions to establish the convergence rates of the adversarial min-max estimator and the corresponding regret bounds. The assumptions are two fold: First, Assumptions (ref) -- (ref) form the foundation for the consistency of the estimator, which is summarized in Lemma (ref). Second, with the definition of a pseudometric and a link condition, Assumption (ref) facilitates the convergence rates and regret bounds.
We first impose some basic assumptions on identification, function spaces, sample criteria, and penalty parameters.
Assumption (ref)(ref) states that the true $\alpha_0$ can be captured by the user-defined space ${\mathcal{H}}$, and the solution of conditional moment restrictions is unique in $({\mathcal{H}},\|\bullet\|_{2,2})$. The uniqueness is typically required in the literature on conditional moment problems ai2003efficient,chen2012estimation. With Assumption (ref)(ref) on the space ${\mathcal{F}}$, we are able to find a good adversary function $f$ for any $\alpha$ in the neighborhood of $\alpha_0$. Assumption (ref)(ref) guarantees that the change of $\alpha$ measured by $W(Z;\alpha)$ can be continuously projected in ${\mathcal{F}}$ space. Assumption (ref)(ref) is also commonly imposed in the literature such that there is no degenerate conditional moment restrictions.
Assumption (ref) is a common assumption in the M-estimation theory van2000asymptotic and holds by implementing the optimization algorithm correctly. Assumption (ref) imposes the conditions on the asymptotic rate of tuning parameters according to the critical radii of user-defined function classes, which are typically required in penalized M-estimation methods. The following lemma on the consistency of the min-max estimator follows directly from Assumptions (ref) -- (ref).
Given the consistency results in Lemma (ref), to obtain a local convergence rate, we can restrict the space ${\mathcal{H}}$ to a neighborhood of $\alpha_0$ defined as
In this restricted space, following chen2012estimation, we define the pseudometric $\|\alpha_1-\alpha_2\|_{ps,\alpha}$ for any $\alpha_1,\alpha_2\in{\mathcal{H}}_{\alpha_0,M_0,\epsilon}$ as
where the pathwise derivative in the direction $\alpha-\alpha_0$ evaluated at $\alpha_0$ is defined as
We further introduce the following assumption for the restricted space ${\mathcal{H}}_{\alpha_0,M_0,\epsilon}$.
Assumption (ref)(ref) and Assumption (ref)(ref) ensure that the pseudometric is well defined in the neighborhood of $\alpha_0$, and Assumption (ref)(ref) restricts local curvature of $\Phi(\alpha)$ at $\alpha_0$, which enables us to attain a fast convergence rate in the sense of $\|\bullet\|_{ps,\alpha_0}$.
In this subsection, we establish general convergence rate of the min-max estimator and the regret bound of learned policy $\widehat{\pi}$, which depend on sample size, complexities of spaces related to ${\mathcal{F}}$, ${\mathcal{H}}$, and the local modulus of continuity. Furthermore, concrete examples and results will be given in Subsection (ref).
To derive the convergence rate in $\|\bullet\|_{2,2}$, we introduce the local modulus of continuity $\omega(\delta, {\mathcal{H}}_M)$ at $\alpha_0$, which is defined as
The local modulus of continuity $\omega(\delta, {\mathcal{H}}_{\alpha_0,M_0,\epsilon})$ enables us to link the local errors quantified by $\|\widehat{\alpha}_n - \alpha_0\|_{ps,\alpha_0}$ and $\|\widehat{\alpha}_n - \alpha_0\|_{2,2}$. For a detailed discussion, see chen2012estimation. We now present a general theorem for the convergence rate in $\|\bullet\|_{ps,\alpha_0}$ and $\|\bullet\|_{2,2}$.
Theorem (ref) states that the convergence rate in pseudometric depends on the critical radii, which measures the complexities, of spaces related to ${\mathcal{F}}$ and ${\mathcal{H}}$. The convergence rate in $\|\bullet\|_{2,2}$ is a direct consequence by applying the local modulus of continuity. Following Theorem (ref), we have the following general regret bound for our estimated policy $\widehat \pi$.
Theorem (ref) shows that bounding the regret of $\widehat{\pi}$ is as easy as bounding $\|\bullet\|_{2,2}$ error of the nuisance functions $\beta_{p,1}$ and $\beta_{p,2}$.
To apply the general convergence rate and regret bound in Theorems (ref) and (ref), we need to compute the upper bound of critical radii $\delta_n$, and local modulus of continuity $\omega(\delta_n,{\mathcal{H}}_{\alpha_0,M_0,\epsilon})$. In this subsection, we focus on the case when ${\mathcal{H}}$ and ${\mathcal{F}}$ are some RKHSs and provide sufficient conditions to bound these terms, which lead to concrete results on convergence rates and regret bounds.
RKHSs endowed with a kernel of polynomial eigenvalue decay rate are commonly used in practice. For example, the $\gamma$-order Sobolev space has a polynomial eigen decay rate. Neutral networks with a ReLU activation function can approximate functions in $H_A^{\gamma}([0,1]^d)$, the $A$-ball in the Sobolev space with an order $\gamma$ on the input space $[0,1]^d$ korostelev2011mathematical.
To obtain the rate of convergence in $\|\bullet\|_{2,2}$, we now quantify the local modulus of continuity for RKHS ${\mathcal{H}}$. By Mercer's theorem, we can decompose $\Delta\alpha = \alpha - \alpha_0\in{\mathcal{H}}_{M_0}$ by the eigen decomposition of kernel $K_{{\mathcal{H}}}$ by $\Delta\alpha = \sum_{j=1}^{\infty} a_j e_j$, where $e_j: {\mathcal{Z}}\rightarrow {\mathbb{R}}^d$ with $\|e_j\|_{2,2}^2=1$ and $\left\langle e_i,e_j \right\rangle_{2,2}=0$. Then $\|\alpha-\alpha_0\|_{2,2}^2 = \sum_{j=1}^{\infty}a_j^2$ and $\|\alpha-\alpha_0\|_{{\mathcal{H}}}^2 = \sum_{j=1}^{\infty} a_j^2/\lambda_j(K_{{\mathcal{H}}}) \leq M_0$. Therefore, $\sum_{j\geq m} a_j^2 \leq \lambda_m(K_{{\mathcal{H}}})M_0$. In addition,
For any positive integer $m$, let
Then we have the following lemma on the local modulus of continuity.
Lemma (ref) allows us to quantify the local modulus of continuity at $\alpha_0$ for the RKHS case by finding an optimal $m^{*}$, which is determined by the decay rate of $\tau_m$. We consider two cases in the following corollary.
Corollary (ref) considers two scenarios of the local modulus of continuity. If $b$ is large for the mild ill-posed case or the severe ill-posed case is considered, the convergence rate of the regret can be much slower. While for the mild ill-posed case with $b\rightarrow0$, we nearly attain the minimax optimal rate, $n^{-\frac{1}{2+1/\min(\gamma_{{\mathcal{H}}},\gamma_{{\mathcal{F}}})}}$, in the classical non-parametric regression stone1982optimal.
In this section, we perform thorough simulation studies to evaluate the numerical performance of the proposed pricing policy learning method, PRINT, in terms of revenue regret. Two benchmark offline pricing policy learning methods are compared with our method under various simulation settings and a dataset from an online auto loan company.
We first generate $(G,X,U)$ to satisfy Assumptions (ref)(ref), (ref), and (ref), in which Assumption (ref)(ref) can be guaranteed by Assumption (ref). Specifically, let $X \sim {\mathcal{N}}((0.25,0.25)^{\top}, \Sigma_x={\mathbf{I}}_2)$. Then we generate $G$ and $U=(U_1,U_2)$ by
According to structural equations (ref) and (ref), we generate price $P$ and revenue $Y$ with previously generated $(G,X,U)$. The coefficient functions
It can be verified that they satisfy Assumption (ref)(ref) and Assumption (ref) and that $\beta_{p,2}(u,x)\leq -c_{p,2}$ for some $c_{p,2}>0$ for all $(u,x)$. Based on structural equations (ref) and (ref), we add two independent noises to generate $Y$ and $P$ respectively, i.e.,
where $\epsilon_y, \epsilon_p\sim \text{Uniform}[-1,1]$, independently.
Under the above setting, an optimal pricing policy is given below, which will be used to calculate the regret.
To control the degree of violation to Assumption (ref)(ref) (IV exclusion restriction), in the function $\beta_g(U,X,G)$, we consider $c_4=1$ for a mild violation and $c_4=5$ for a severe violation. To control the strength of the invalid IV $G$, in $\alpha_g(U,X,G)$, we set $c_7=1$ for a weak IV and $c_7=5$ for a strong IV. The values of other parameters are given in Appendix (ref).
We evaluate the learned policy 100 times by the Monte Carlo method on a noise-free testing dataset of size 10,000. Figure (ref) shows the box plots of the regrets of revenue (${\mathcal{V}}(\pi^*) - {\mathcal{V}}(\widehat{\pi})$) of learned policies by different methods from 100 replicates of simulation. For all combinations of IV strength and the violation of exclusion restriction, the policies learned by the PRINT outperform the other two benchmark methods for both sample sizes $n=1,000$ and $2,000$ and can achieve better performance with a larger sample size $n=2,000$. When the IV is strongly relevant the price, the policies learned by the regression could achieve small regrets. However, when the IV is weakly relevant to the price, the performance of policies learned by regression is unstable. Finally, the overall performance of linear policies learned by kallus2018policy is not stable, partially due to that the inverse of the generalized propensity scores is hard to estimate, and the confounding bias caused by the unmeasured $U$.
In this subsection, we study the numerical performance of the pricing strategy of personalized loans for an anonymous US auto lending company. We compare our method by kallus2018policy, a regression-based method, and the historical decision made by the company. A major difference between the real data application and the previous simulation study is that the structural equations (ref) and (ref) are potentially misspecified in the real data.
We obtain the dataset CPRM-12-001: On-Line Auto Lending from the Center for Pricing and Revenue Management at Columbia University\footnote{\url{https://www8.gsb.columbia.edu/cprm/research/datasets}}. It records the online auto loan applications received by the company from Jul. 2002 to Nov. 2004. For each approved application, the requested term, loan amount, annual percentage rate (APR), monthly London interbank offered rate (LIBOR), whether contracted or not, and some personal information (e.g., FICO score) are recorded. For detailed descriptions, we refer the readers to phillips2015effectiveness and ban2021personalized.
For the pricing of the online auto loan company, we adopt the price defined in ban2021personalized, which is the net present value of future payments less the loan amount, i.e.,
We set the feasible price range to be $[\$0, \$40,000]$.
To evaluate the pricing policy, it is necessary to construct a generative model since the revenue depends on the price selected by the policy, which is not available in the dataset. This is in contrast to supervised learning, where a testing dataset can be used to evaluate prediction accuracy. Since the outcomes of whether contracted or not are binary (accept/reject), instead of a linear demand, we adopt a logistic demand model used by ban2021personalized for generating the demand. Therefore the true expected revenue is not a quadratic function of the price, conditioning on other factors. While this causes a model mis-specification for our method, we indeed find that our method works well and is robust to such a mis-specified revenue model. In addition, this serves as a complement to the previous simulation study, where we studied a correctly specified model. In particular, for a given feature vector $x$ (including the constant term) and a price $p$, the probability of accepting the contract, i.e., the expected demand, is modeled by $ \frac{1}{1+ \exp \{ -\alpha^{\top} x - \beta^{\top}x\times p\}}. $ Then the expected revenue is $p/[1+ \exp \{ -\alpha^{\top} x - \beta^{\top}x\times p \}]$. The population demand model, which will be used to evaluate the revenue, is obtained by fitting a logistic regression with $\ell^2$-penalty with all records. We select the penalty parameter by 5-fold cross-validation.
It is worth noting that even a competitor's rate is available in this dataset. In practice, however, we actually do not have reliable data to represent the competition in the auto lending industry during the analyzed period of time phillips2015effectiveness. Therefore, we intentionally treat the competitor's rate as an unmeasured confounding and exclude it from the feature vector. Further, given the loan amount, term, and LIBOR, the price of a loan can be purely determined by the monthly payment. Therefore, to ensure that Assumption (ref) (ref) holds, we also exclude the monthly payment from the feature vector.
For PRINT, we choose the APR (annual percentage rate) for a loan as the instrumental variable, which was used as the instrument for dealing with the endogeneity of the price by blundell1992credit. Meanwhile, the loan rate is continuous and has a strong direct effect on the continuous price, which indicates that it could be used as a strong relevant IV. However, using the loan rate as a valid IV is questionable, since it may have a direct effect on the eventual revenue, hence breaking the IV exclusion restriction. However, we can safely use the loan APR as the invalid IV for our method since the IV exclusion restriction has been relaxed.
For comparison, in addition to the two benchmark methods compared in Section (ref), we also consider the company's actual pricing policy and the optimal policy by maximizing the revenue according to the learned population demand model. We use the first 60,000 records (ordered by application time) of the dataset as the training data to learn the policies for PRINT and two benchmark methods. Then we apply those learned policies to the testing dataset of the rest 148,084 records and calculate the revenues by the learned demand model.
Table (ref) summarizes the expected revenue for the pricing policy learned by the PRINT against the benchmark policies, optimal policy, and policy used by the firm. The firm's historical policy could attain 77.3% of the revenue if the optimal policy were used. This satisfactory revenue is reasonable since the firm may have external information at the pricing time that is not available in the offline records. The expected revenues obtained by the policies learned by direct regression and kallus2018policy do not reach the firm's historical revenue, partially because of the unobserved confounding issue and insufficient offline data. In contrast, the policy learned by PRINT attains about 82.3% of the expected revenue of the optimal policy and improves the firm's revenue by 6.5%.
In Figure (ref), we provide the density plots of the prices by the five pricing policies on the testing dataset. It is clear that the distribution of prices suggested by PRINT is the closest to the oracle optimal policy.
To study the interpretability of the policy learned by PRINT, in Figure (ref), we provide the partial dependence plots of the learned policy on the four most important features. First, as shown in Figure (ref)(a) and (b), the price increases with LIBOR and term, which agrees with the definition of the price in (ref). Second, Figure (ref)(c) shows that for customers with higher FICO scores, we suggest lower prices as those customers have low risks of bad debts. Finally, as shown in Figure (ref)(d), our policy tends to set higher prices for the applications with longer processing times, this is likely because those applications have potential risks that require longer review processes.
In this paper, we study offline personalized pricing under endogeneity by leveraging an instrumental variable. The key challenges are (a) identification of the heterogeneous effect of continuous price on revenue under unmeasured confounding; (b) a possibly invalid IV that may violate the exclusion restriction; (c) solving conditional moment restrictions of generalized residual functions. For (a) and (b), we generalized the identification results in causal inference literature on relaxing exclusion restriction for discrete treatment to continuous treatment. For (c), we develop an adversarial min-max algorithm for learning the optimal pricing strategy. Theoretically, we established the consistency and the convergence rate of the proposed policy learning algorithm. For future work, it is interesting to extend our work to learning a multi-stage policy strategy with offline data under endogeneity with invalid IVs.