EconBase
← Back to paper

Welfare Analysis via Marginal Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

31,260 characters · 13 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Welfare Analysis via Marginal Treatment Effects

abstractConsider a causal structure with endogeneity (i.e., unobserved confoundedness) in empirical data, where an instrumental variable is available. In this setting, we show that the mean social welfare function can be identified and represented via the marginal treatment effect bjorklund/moffitt:1987 as the operator kernel. This representation result can be applied to a variety of statistical decision rules for treatment choice, including plug-in rules, Bayes rules, and empirical welfare maximization (EWM) rules as in hirano2020. Focusing on the application to the EWM framework of kitagawa2018should, we provide convergence rates of the worst case average welfare loss (regret) in the spirit of manski:2004. \begin{description} • {\bf Keywords:} empirical welfare maximization (EWM), endogeneity, heterogeneity, marginal treatment effects (MTE), statistical decision rules, treatment choice • {\bf JEL Codes:} C14, C21 \end{description}

Introduction

One of the most important goals of empirical economic research is to advise policy makers on how to assign heterogeneous individuals to a treatment under consideration subject to budgetary, legal, and ethical constraints, based on evidence from empirical data. To this goal, it is crucial to identify a social welfare function in observational data settings. For many observational data sets used by empirical researchers, treatments are likely to be endogenously selected by rational individuals, rather than randomly assigned. Furthermore, the effects of these treatments are often heterogeneous across individuals, even after controlling for their observable attributes. In this light, we propose a novel method of identifying the mean social welfare function in the presence of unobserved heterogeneity in treatment effects while accounting for endogenous treatment selection in empirical data.

It is well known today that the marginal treatment effects bjorklund/moffitt:1987 measure heterogeneous treatment effects, and the MTE can be identified with an instrumental variable under endogenous treatment selection heckman/vytlacil:2001,heckman/vytlacil:2005,heckman/vytlacil:2007. Hence, it is a natural idea to use the MTE as a building block for the identification of a social welfare function in the presence of unobserved heterogeneity and endogeneity. In this paper, we show that the mean social welfare function can be indeed identified and represented via the MTE as the operator kernel. Since the identification and estimation of the MTE have been well established in the existing literature heckman/vytlacil:2001,heckman/vytlacil:2005,heckman/vytlacil:2007,carneiro/lee:2009,carneiro/heckman/vytlacil:2010, our result thus paves the way for these existing theories and methods of MTE to be directly applied to welfare analysis.

Once the mean social welfare function has been identified via the MTE, we can apply it to a variety of policy makers' statistical decision problems of treatment choice, including those based on plug-in rules, Bayes rules, and empirical welfare maximization (EWM) rules hirano2020. Focusing on the EWM rules in particular, we can take advantage of the technology developed by kitagawa2018should to analyze properties of the EWM method in the spirit of manski:2004. Specifically, under both the heterogeneity and endogeneity, we can derive convergence rates of the worst case average welfare loss (regret) from the maximum empirical welfare. As such, our result contributes to the literature by extending the scope of applicability of the EWM framework of kitagawa2018should, that is originally based on the assumption of selection on observables or unconfoundedness, to the framework that now allows for unobserved confoundedness or endogeneity.

A recent paper by athey/wager:2020 considers an endogeneity problem in the context of the EWM, and proposes a theory based on doubly robust estimators of average treatment effects. The result that we propose in this paper neither nests nor is nested by that of athey/wager:2020 -- these two papers play rather complementary roles. On the one hand, athey/wager:2020 assume homogeneous treatment effects, which implies a constant MTE, while our framework can allow for unobserved heterogeneity in treatment effects. On the other hand, the applicability of our proposed method hinges on the identification of the MTE, while athey/wager:2020 do not need to identify the MTE for their objective. In other words, our framework accommodates unobserved heterogeneity at the expense of assuming that the MTE is identified. This tradeoff illustrates the complementarity between our result and the result developed by athey/wager:2020.

Another recent paper by byambadalai:2020 also considers an endogeneity problem in the context of counterfactual welfare comparisons, and is thus closely related to this paper. These two papers again play complementary roles. On the one hand, byambadalai:2020 develops the partial identification and focuses on welfare gains and losses as parameters of interest, while we develop the point identification of the welfare function that is applicable to various statistical decision rules as well as welfare comparisons. On the other hand, byambadalai:2020 imposes weak assumptions and in particular does not need to assume to identify the MTE. This tradeoff illustrates the complementarity between our result and the result developed by byambadalai:2020.

{\bf Relation to the Literature:} This paper aims to contribute to the literature on statistical decisions in econometrics -- see the recent survey by hirano2020 for a comprehensive review of this subject. In particular, we focus on an application of our representation theorem to bounding the worst case average welfare loss in the spirit of manski:2004 with the recent technology developed by kitagawa2018should, as mentioned above. Besides, this paper is also related to the broad literature on policy choices and welfare analysis including manski:2004, manski:2004,manski:2009, dehejia2005program, schlag2007eleven, hirano2009asymptotics,hirano2020, stoye2009minimax,stoye2012minimax, chamberlain2011bayesian, bhattacharya2012inferring, tetenov2012statistical, armstrong2015inference, kasy:2016, mbakop2016model, kock2017optimal, kitagawa2018should,kitagawa2019equality, rai2018statistical, sakaguchi:2019 viviano2019policy, athey/wager:2020, byambadalai:2020, han2020comment, and sun:2020. In particular, our result complements those of athey/wager:2020 and byambadalai:2020, as mentioned above. Also closely related is the literature on the MTE bjorklund/moffitt:1987 and its identification and estimation, including heckman/vytlacil:2001,heckman/vytlacil:2005,heckman/vytlacil:2007, carneiro/lee:2009, carneiro/heckman/vytlacil:2010, brinch/mogstad/wiswall:2017, lee2018identifying, and mogstad/santos/torgovitsky:2017 among many others. The applicability of our representation theorem relies on the identification of the MTE from this literature. Finally, this paper also complements the literature on policy relevant treatment effects heckman/vytlacil:2001,heckman/vytlacil:2005,heckman/vytlacil:2007,brinch/mogstad/wiswall:2017,carneiro/lokshin/umapathi:2017,mogstad/santos/torgovitsky:2017,sasaki2018estimation which develops methods of identification, estimation and inference for average welfare gains under counterfactual policies based on the MTE.

{\bf Organization:} The rest of this paper is organized as follows. Section (ref) introduces the model. Section (ref) presents the main result of representing the mean social welfare via the marginal treatment effects. Section (ref) introduces applications to three statistical decision rules. Section (ref) demonstrates the use of the representation result in the empirical welfare analysis. Section (ref) concludes. The appendix contains mathematical proofs.

The model

Consider the model \begingroup \allowdisplaybreaks

align[align omitted — 98 chars of source]

\endgroup where $Y$ denotes an observed outcome variable, $D$ denotes an observed binary treatment variable, $Z$ denotes a vector of observed exogenous variable, $Y_0$ and $Y_1$ denote unobserved potential outcomes under no treatment and under treatment, respectively, and $\tilde{U}$ denotes an unobserved factor of the treatment selection. The first equation (ref) models the outcome production through the potential outcome framework, and the second equation (ref) models the treatment selection via a threshold-crossing model. The function $\tilde{\nu}$ in this threshold-crossing treatment assignment model (ref) is nonparametric and is unknown to the econometrician.

This model allows for endogeneity (unobserved confoundedness) in the sense that $(Y_0,Y_1)$ and $\tilde{U}$ may be statistically dependent even conditionally on $Z$. For the purpose of identification, therefore, we require the vector $Z$ to consist of excluded exogenous variables (i.e., excluded instruments) as well as included exogenous variables, as formally stated in Assumption (ref) below. The next assumption are standard in the recent literature on marginal treatment effects brinch/mogstad/wiswall:2017,mogstad/santos/torgovitsky:2017.

assumption[Model Restrictions] The random vector $Z$ can be written as $(Z_0',X')'$, where \begin{enumerate} • $\tilde{U}$ and $Z_0$ are independent given $X$; • $E[Y_d\mid Z,\tilde{U}]=E[Y_d\mid X,\tilde{U}]$ and $E[Y_d^2]<\infty$; and • $\tilde{U}$ is continuously distributed with a convex support conditional on $X$. \end{enumerate}

Part (i) concerns solely about the treatment assignment model (ref), and this is the only independence assumption to be imposed on the model, implying that we can allow for an arbitrary statistical dependence between the potential outcomes $(Y_0,Y_1)$ and $\tilde U$, even conditionally on $Z$. Part (ii) states the exclusion restriction of the random sub-vector $Z_0$ of $Z$, and bounded second moments of the potential outcomes $(Y_0,Y_1)$. Part (iii) rules out point masses and holes in the conditional distribution of $\tilde{U}$ given $X$.

For ease of analysis, by following the literature on the marginal treatment effects, we apply normalizing transformations, $U\equiv F_{\tilde{U}\mid X}(\tilde{U})$ and ${\nu}(Z)\equiv F_{\tilde{U}\mid X}(\tilde{\nu}(Z))$, in the threshold crossing model (ref). The following lemma confirms convenient properties to be used throughout the rest of the paper, as a result of these normalizing transformations under Assumption (ref).

lemma[Normalization] Suppose that Assumption (ref) (i) and (iii) hold. Then (i) $D=1\{\nu(Z)-U\geq 0\}$, and (ii) $U$ is distributed uniformly over $[0,1]$ conditional on $Z$.

A proof of this lemma is provided in Appendix (ref). Consequently, we can rewrite the threshold-crossing treatment selection model (ref) without loss of generality as

equation[equation omitted — 98 chars of source]

We will hereafter substitute the model model (ref) for the original model (ref) by virtue of Assumption (ref).

The main result

In this section, we show that the social welfare function can be identified and represented via the marginal treatment effects as the operator kernel. To this end, we first introduce and define the two key ingredients of this result, namely the social welfare function and the marginal treatment effects.

A policy maker assigns individuals with certain observed attributes $Z$ to a treatment $D=1$. Thus, a treatment assignment rule is represented by a decision set $G \subset \mathcal{Z}$, where $\mathcal{Z}$ is a set of values that $Z$ may take. Specifically, the decision set $G$ represents the policy in which individuals with $Z \in G$ are assigned to a treatment $D=1$ while those with $Z \not\in G$ are not. Let $\mathcal{G}$ denote a collection all the decision sets $G$ under consideration subject to the policy makers' constraints. With these notations, the social welfare function $W: \mathcal{G} \rightarrow \mathbb{R}$ is defined by $$ W(G)=E[1\{Z\in G\}Y_1+1\{Z\notin G\}Y_0]. $$ Next, recall from Assumption (ref) that $X$ is the included sub-vector of the random vector $Z$ of exogenous variables that affects the treatment assignment, and also recall from Lemma (ref) or Equation (ref) that $U$ is the normalized unobserved factor of the treatment selection. The marginal treatment effect bjorklund/moffitt:1987 is defined by $$ MTE(u,x)=E[Y_1-Y_0\mid U=u,X=x]. $$ With these definitions of the social welfare function and the marginal treatment effect, we now state the following theorem as the main result of this paper.

theorem[Representation] Under Assumption (ref), one has \begin{equation} W(G)=E[Y_0]+E\left[1\{Z\in G\}\int_0^1MTE(u,X)du\right] \qquadfor every $G \in \mathcal{G}$. \end{equation}
proofSince $Y_1$ and $Y_0$ are integrable under Assumption (ref) (ii), we have \begingroup \allowdisplaybreaks \begin{align*} W(G) =& E[1\{Z\in G\}Y_1+1\{Z\notin G\}Y_0]\\ =& E[1\{Z\in G\}(Y_1-Y_0)]+E[Y_0]. \end{align*} \endgroup Now, the statement of this theorem follows from \begingroup \allowdisplaybreaks \begin{align*} E[1\{Z\in G\}(Y_1-Y_0)] =& E[1\{Z\in G\}E[Y_1-Y_0\mid Z,U]]\\ =& E[1\{Z\in G\}MTE(U,X)]\\ =& E[1\{Z\in G\}E[MTE(U,X)\mid Z]]\\ =& E\left[1\{Z\in G\}\int_0^1MTE(u,X)du\right], \end{align*} \endgroup where the first equality follows from the law of iterated expectations, the second equality follows from Assumption (ref) (ii) and the definition of $MTE$, the third equality follows from another application of the law of iterated expectations, and the fourth equality follows from Lemma (ref) under Assumption (ref) (i) and (iii).

The representation (ref) of the mean social welfare via $MTE(u,x)$ is the key result of this paper. Because there is an existing literature on identification and estimation for the MTE heckman/vytlacil:2001,heckman/vytlacil:2005,heckman/vytlacil:2007,carneiro/lee:2009,carneiro/heckman/vytlacil:2010, our result (ref) paves the way for empirical welfare analysis under the potential endogeneity or unobserved confoundedness based on the existing identification and estimation methods of the MTE. We present a few examples of such applications in Sections (ref) and (ref).

Finally, we remark on relations to and differences from kitagawa2018should, who use the representation $$ W(G) = E[Y_0] + E\left[1\{Z \in G\} \tau(X)\right], $$ where $\tau(x) \equiv E[Y_1-Y_0|X=x]$. Our representation (ref) is closely related to this representation. Under the unconfoundedness assumption, kitagawa2018should use the identification of $\tau(x)$ by $E[Y|D=1,X=x]-E[Y|D=0,X=x]$. On the other hand, under the unobserved confoundedness in our setup, the corresponding operator kernel $\tau(x)$ is not identified by $E[Y|D=1,X=x]-E[Y|D=0,X=x]$ in general. Instead, we propose to take advantage of the identification and estimation of $MTE$ from the literature on the marginal treatment effects.

Applications to statistical decision rules

Once we obtain the representation (ref) of the mean social welfare via $MTE(u,x)$, we may apply it to a variety of policy makers' statistical decision problems for treatment choice. In this section, following hirano2020, we introduce applications to the three popular statistical decision rules: 1. plug-in rules, 2. Bayes rules, and 3. empirical welfare maximization rules. In Section (ref), we discuss the empirical welfare maximization rules in further details based on recent technologies.

Plug-in rules

Suppose that the distribution of $Z$ is parametrized by $\delta$ and we have an estimator $(\hat\delta,\widehat{MTE})$ for $(\delta,MTE)$. In this case, we can estimate the maximizer for the population social welfare by maximizing $$ \int 1\{z\in G\}\int_0^1\widehat{MTE}(u,x)dud\mu_{\hat\delta}(z) $$ over $G\in\mathcal{G}$, where $\mu_{\delta}(\cdot)$ is the probability measure of $Z$ indexed by $\delta$ and we may ignore the term $E[Y_0]$ in (ref) since it does not affect the maximization problem over $G\in\mathcal{G}$. Note that $(\hat\delta,\widehat{MTE})$ is an estimator, and so a maximizer of the above objective function is a statistical decision rule, i.e., it is a function of the observed data.

Bayes rules

Suppose that the distribution of $Z$ is parametrized by $\delta$ and $MTE$ is parametrized by $\eta$. In addition, suppose that we have a prior probability measure, denoted by $\pi_{\text{prior}}$, of $(\delta,\eta)$, and we can construct the posterior distribution, denoted by $\pi_{\text{posterior}}$, of $(\delta,\eta)$ by Bayesian updating. Given the posterior probability measure of $(\delta,\eta)$ and our representation (ref), we can construct the Bayes welfare $$ \int \left(\int 1\{z\in G\}\int_0^1{MTE}_\eta(u,x)dud\mu_\delta(z)\right)d\pi_{\text{posterior}}(\delta,\eta). $$ The Bayes rule is the maximizer of the Bayes welfare over $G \in \mathcal{G}$.

Empirical welfare maximization rules

The empirical welfare maximization rule uses the empirical distribution of $Z$ and an estimator $\widehat{MTE}$ for $MTE$. Namely, with a random sample $\{Z_1,\ldots,Z_n\}$ of size $n$, we can define the empirical welfare by $$ E_n\left[1\{Z\in G\}\int_0^1\widehat{MTE}(u,X)du\right], $$ where $E_n$ denotes the sample average operator, i.e., $E_n f(Z) = n^{-1} \sum_{i=1}^n f(Z_i)$ for any measurable function $f$. The empirical welfare maximization rule selects the maximizer of this empirical welfare over $G \in \mathcal{G}$. The following section presents more detailed analyses of the asymptotic properties of the maximum of this empirical welfare relative the population mean welfare under the oracle action.

Applications to empirical welfare maximization

We demonstrate applications of the representation result (ref) to empirical welfare maximization in this section. For the purpose of exposition of the core idea, we first focus on the case where the mapping $(u,x)\mapsto MTE(u,x)$ is known by a researcher in Section (ref). We then present the case where the mapping $(u,x)\mapsto MTE(u,x)$ is unknown by a researcher and thus needs to be estimated in Section (ref).

Empirical welfare maximization with known MTE

In this section, we assume that we know the mapping $(u,x)\mapsto MTE(u,x)$. The empirical welfare maximizer in this setting is given by $$ \hat{G}_{EWM}\in\arg\max_{G\in\mathcal{G}}E_n\left[1\{Z\in G\}\int_0^1MTE(u,X)du\right]. $$ We present a uniform asymptotic analysis of the maximum empirical welfare $W(\hat{G}_{EWM})$, relative to the population mean social welfare under the oracle action, denoted by $$ W_{\mathcal{G}}=\sup_{G\in\mathcal{G}}W(G). $$ To this end, consider the following assumption.

assumption(i) $|\int_0^1MTE(u,X)du|\leq\bar{M} < \infty$ a.s. for the class $\mathcal{P}(\bar{M})$ of distributions of $(Y_0,Y_1,D,Z)$. (ii) $\mathcal{G}$ has a finite VC-dimension $v<\infty$ and is countable.

The above assumption is a modification of Assumption 2.1 (BO)-(VC) in kitagawa2018should tailored to our framework with the marginal treatment effects. Assumption (ref) (i) requires a bounded integral of the marginal treatment effect function. As a sufficient condition, it holds when the outcome variable is bounded by some constant which is naturally satisfied in some applications. Assumption (ref) (ii) restricts the complexity of the class of treatment functions $Z\mapsto 1\{Z\in G\}$. As a sufficient condition, when $X$ has a finite support, this assumption will automatically hold where $v$ is the cardinality of the power set for the support of $X$. kitagawa2018should collects a few examples of $\mathcal{G}$ with finite VC-dimensions.

The following corollary provides a convergence rate of the worst case average welfare loss (regret) by the empirical welfare maximization.

corollaryUnder Assumptions (ref) and (ref), one has $$ \sup_{P\in\mathcal{P}(\bar{M})}E_{P^n}\left[W_{\mathcal{G}}-W(\hat{G}_{EWM})\right]\leq 2C_1\bar{M}\sqrt{\frac{v}{n}}, $$ where $C_1$ is a universal constant.

A proof is provided in Appendix (ref). Corollary (ref) implies that no treatment assignment rule based on empirical data will achieve a minimax rate that is faster than $n^{-1/2}$, and shows that $\hat{G}_{EWM}$ is the minimax rate optimal over $\mathcal{P}(\bar{M})$. This corollary extends and is a counterpart of Theorem 2.1 of kitagawa2018should.

Empirical welfare maximization with unknown MTE

In this section, we consider the case where the mapping $(u,x)\mapsto MTE(u,x)$ is unknown by a researcher and thus needs to be estimated from empirical data. The empirical welfare maximizer in this setting is given by $$ \hat{G}_{hybrid}\in\arg\max_{G\in\mathcal{G}}E_n\left[1\{Z\in G\}\int_0^1\widehat{MTE}(u,X)du\right], $$ where $\widehat{MTE}$ is an estimator of $MTE$. The existing literature on marginal treatment effects provides a list of alternative estimators $\widehat{MTE}$ of $MTE$. We therefore first provide a general sufficient condition in terms of $\widehat{MTE}$ that accommodate a wide range of possible estimators in Section (ref). This will be followed up by a specific estimator $\widehat{MTE}$ with lower level primitive conditions tailored to it in Section (ref).

A sufficient condition

A general high-level assumption for an estimator $\widehat{MTE}$ of $MTE$ is that it entails a uniform convergence rate $\psi_n^{-1}$ in the mean absolute value of $\int_0^1 \widehat{MTE}(u,X)du$ for $\int_0^1 MTE(u,X)du$, as formally stated below.

assumptionFor a class of data generating processes $\mathcal{P}_m$, there exists a sequence $\psi_n\rightarrow\infty$ such that $$ \limsup_{n\rightarrow\infty}\sup_{P\in\mathcal{P}_m}\psi_nE_{P^n}\left[E_n\left[\left|\int_0^1(\widehat{MTE}(u,X)-{MTE}(u,X))du\right|\right]\right]<\infty. $$

This sufficient condition leads to the rate, $\psi_n^{-1} \vee n^{-1/2}$, of convergence for the worst case average welfare loss (regret), as formally stated as a corollary below.

corollaryUnder Assumptions (ref), (ref), and (ref), $$ \sup_{P\in\mathcal{P}_m\cap\mathcal{P}(\bar{M})}E_{P^n}\left[W_{\mathcal{G}}-W(\hat{G}_{hybrid})\right]=O(\psi_n^{-1}\vee n^{-1/2}). $$

A proof of this corollary is provided in Appendix (ref). It serves as a counterpart of Theorem 2.5 of kitagawa2018should. Different estimators $\widehat{MTE}$ of $MTE$ in general entail different convergence rates $\psi_n^{-1}$. We next present a concrete estimator with lower-level sufficient conditions for the high-level condition in Assumption (ref).

Parametric estimation for $MTE(u,x)$

By heckman/vytlacil:1999,heckman/vytlacil:2001,heckman/vytlacil:2005, the MTE can be identified from data via the equality $$ MTE(u,x)=\frac{E[Y\mid \nu(Z)=u,X=x]}{\partial u}\equiv LIV(u,x). $$ In this section, we consider the parametric regression function for $E[Y\mid \nu(Z)=u,X=x]$ by $$ E[Y\mid \nu(Z)=u,X=x]=x'\beta_0+x'(\beta_1-\beta_0)u+\sum_{k=2}^K\alpha_ku^k $$ following cornelissen/dustmann/raute/schonberg:2016 on the survey of the marginal treatment effects for labor economists. Let $\hat\nu(Z)$ denote some estimator of the propensity score $\nu(Z)$, and define \begingroup \allowdisplaybreaks

align*[align* omitted — 239 chars of source]

\endgroup

Let $\hat\theta=(\hat\beta_0,\hat\beta_1,\hat\alpha_2,\ldots,\hat\alpha_K)$ be the OLS estimator for $\theta$ by regressing $Y$ on $\hat{\mathcal{X}}$, that is, $$ \hat\theta=E_n\left[\hat{\mathcal{X}}\hat{\mathcal{X}}'\right]^{-1}E_n\left[\hat{\mathcal{X}}Y\right]. $$ Then, $MTE$ can be simply estimated by the following linear functional of $\hat\theta$. $$ \widehat{MTE}(u,x)=x'(\hat\beta_1-\hat\beta_0)+\sum_{k=2}^Kk\hat\alpha_ku^{k-1}. $$ Therefore, the operator kernel of in our representation (ref) can be estimated by the simple linear expression $$ \int_0^1\widehat{MTE}(u,x)du=x'(\hat\beta_1-\hat\beta_0)+\sum_{k=2}^K\hat\alpha_k. $$ For this concrete estimator, we provide a set of lower-level conditions in the proposition below that guarantee the aforementioned high-level condition in Assumption (ref) to be satisfied.

propositionLet $C$ and $c$ be positive constants, and $\psi_n$ be a sequence with $\psi_n\geq n^{1/2}$. Suppose that the parameter space for $\theta$ is compact so that for sufficiently large $n$ \begin{equation} \|\hat\theta\|+\|\theta\|\leq C almost surely. \end{equation} Furthermore, suppose that $\mathcal{P}_m$ is a class of data generating processes such that \begingroup \allowdisplaybreaks \begin{align} \limsup_{n\rightarrow\infty}\sup_{P\in\mathcal{P}_m}\psi_nE_{P^n}\left[\max_{i=1,\ldots,n}|\hat\nu(Z_i)-\nu(Z_i)|^2\right]^{1/2}<\infty,& \\ \max\{E\left[\|X\|^4\right],E\left[|Y|^4\right]\}<C,& \qquadand \\ \lambda_{\min}\left(E\left[\mathcal{X}\mathcal{X}'\right]\right)\geq c.& \end{align} \endgroup Then, Assumption (ref) is satisfied.

A proof is provided in Appendix (ref). Except for (ref), all the conditions stated in Proposition (ref) can be considered as regularity conditions. The condition in (ref) requires a convergence rate for an estimator $\hat\nu(z)$ of the propensity score $\nu(z)$ uniformly over the data generating processes. This condition can be checked with specific propensity score estimators. For example, we can use the local polynomial estimator $\hat\nu(z)$ for $\nu(z)$, for which kitagawa2018should derive a uniform convergence rate. Specifically, the convergence in (ref) follows directly from their Lemma E.4 (ii). For another example, one could consider a linear propensity score model $\hat\nu(z) = p(z)' \hat\gamma$ and its least squares estimator $\hat\nu(z) = p(z)' \hat\gamma$. In this case, (ref) can be satisfied with $\psi_n = \sqrt{n}$, so that the convergence rate for the worst case average welfare loss (regret) in (ref) holds with the parametric root $n$ rate.

Conclusion

An important research goal for empirical economists is to provide policy makers with guidance on how heterogeneous individuals can be assigned to a treatment under consideration based on evidence from empirical data. To this goal, it is essential to identify a social welfare function from observational data. For many observational data sets used in empirical economic research, treatments are likely to be endogenously selected by rational agents. Furthermore, the effects of these treatments are often heterogeneous even after controlling for observed attributes. In this light, given the abilities of the marginal treatment effects to measure heterogeneous treatment effects, we propose the usage of the marginal treatment effects for identifying the mean social welfare function in the presence of unobserved heterogeneity in treatment effects while accounting for endogenous treatment selection in the empirical data. Our main result, Theorem (ref), establishes that the mean social welfare can be represented via the marginal treatment effects as the operator kernel. We introduce applications of this main result to a few of policy makers' statistical decision problems, such as the plug-in rules, the Bayes rules, and in particular the empirical welfare maximization rule. Focusing on the last application, we derive convergence rates of the worst case average welfare loss (regret) from the maximum empirical welfare under alternative scenarios. The proposed representation in Theorem (ref) can be beneficial as it allows the machinery developed in the existing literature on the marginal treatment effects to be directly applicable to the variety of empirical welfare analysis.

{\singlespacing }