EconBase
← Back to paper

Personalized Subsidy Rules

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,751 characters · 18 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Personalized Subsidy Rules

\thispagestyle{empty}

abstract\onehalfspacing Subsidies are commonly used to encourage behaviors that can lead to short- or long-term benefits. Typical examples include subsidized job training programs and provisions of preventive health products, in which both behavioral responses and associated gains can exhibit heterogeneity. This study uses the marginal treatment effect (MTE) framework to study personalized assignments of subsidies based on individual characteristics. First, we derive the optimality condition for a welfare-maximizing subsidy rule by showing that the welfare can be represented as a function of the MTE. Next, we show that subsidies generally result in better welfare than directly mandating the encouraged behavior because subsidy rules implicitly target individuals through unobserved heterogeneity in the behavioral response. When there is positive selection, that is, when individuals with higher returns are more likely to select the encouraged behavior, the optimal subsidy rule achieves the first-best welfare, which is the optimal welfare if a policy-maker can observe individuals' private information. We then provide methods to (partially) identify the optimal subsidy rule when the MTE is identified and unidentified. Particularly, positive selection allows for the point identification of the optimal subsidy rule even when the MTE curve is not. As an empirical application, we study the optimal wage subsidy using the experimental data from the Jordan New Opportunities for Women pilot study. {\bf Keywords:} Heterogeneous Treatment Effects, Marginal Treatment Effect, Partial Identification, Point Identification, Positive Selection, Shape Restrictions, Welfare Maximization.

Introduction

Governments worldwide have offered various subsidies, such as tuition subsidies for college education, price subsidies for preventive health products, and childcare subsidies for certified daycare services, to encourage self-serving and socially beneficial behaviors among households and individuals. Owing to its relevance and prevalence, many studies across fields in economics have contributed to the evaluation of subsidy programs. For example, in labor economics, empirical works have used exogenous variations in the choice of education to estimate the subsidy on return to tuition ichimura2002semiparametric, carneiro2011estimating. In development economics, researchers have examined the impact of subsidies on the take-up of insecticide-treated bed nets cohen2010free. These studies demonstrate the importance of informing the policy-maker on the optimal allocation of subsidies.

In this paper, we examine subsidy rules that provide personalized subsidies to maximize the targeted population's welfare. A subsidy rule, also called (subsidy-based) policy, is defined as an assignment rule that maps individuals to the amount of subsidy based on their observable characteristics. As the subsidy's effect on both the take-up and welfare outcomes could vary across individuals, an efficient subsidy scheme should consider the welfare effect, behavioral response, and cost of subsidies altogether. Accordingly, an ideal allocation of subsidies increases the take-up among those who benefit the most from it while avoiding those unlikely to benefit from the subsidy. As an advantage, our approach allows for flexible specifications of the cost functions, which can depend on the take-up, individual characteristics, and the amount of subsidy.

We adopt the marginal treatment effect (MTE) framework heckman2005structural to study the subsidy problem. In our setting, the subsidy is an instrumental variable in the following sense. First, the subsidy only affects the take-up of the subsidized behavior (the treatment) but is otherwise excluded from the outcome of interest. In case of tuition, we are effectively assuming that the subsidy affects a student's future income through education received in school. Second, the subsidy is assumed to be randomly assigned (conditional on the covariates) to the population from which our data are sampled, which holds true for any experiment conducted using subsidies. The exclusion restriction and exogeneity of the subsidy help in causally identifying the treatment effect and the individuals' treatment selection process, which in turn helps in identifying the subsidy rule that maximizes the counterfactual welfare by allocating subsidies based on observed heterogeneity.

The MTE framework is appropriate for studying the subsidy problem because of its straightforward yet flexible method for modeling the treatment effects and the treatment selection process. Explicit modeling of the subsidy as an instrument in the treatment selection process helps improve both the interpretability of the results and the practicability of our procedure. Structural assumptions backed by economic theory can be easily incorporated into the model to facilitate identification and estimation of the optimal subsidy rule. As an example, when individuals with higher returns are more likely to select the treatment is the case of positive selection. This can be modeled by assuming that the MTE is decreasing. Our study shows that there are numerous implications of a monotonic MTE in the optimal subsidy problem.

The subsidy problem is analyzed by first characterizing the welfare of subsidy rules using the MTE curve. The MTE can serve as the building block for other treatment parameters such as average treatment effect (ATE) and policy-relevant treatment effect (PRTE) heckman2005structural after suitable reweighting. Consistent with these results, we show that the welfare of a given subsidy rule can also be expressed as a function of the MTE. The welfare characterization can be used to derive necessary conditions for a subsidy rule to be optimal. We show that the necessary condition becomes sufficient when there is positive selection.

Subsidy-based policies are popular because policies that directly mandate the treatment, termed direct policies, are not always viable to the policy-maker owing to practical issues. This study shows that assigning subsidies also achieves higher welfare than directly mandating the treatment, which justifies the subsidy rules. Subsidy rules enjoy this property because, as individuals make treatment choices, subsidy-based policies implicitly target individuals based on their unobserved (to the policy-maker) heterogeneity. For example, individuals make schooling decisions based on any likely private information regarding their returns to higher education. Consequentially, different subsidies may draw students with different returns. If the selection into education is positive, smaller subsidy amounts tend to draw students with higher returns when all other things are equal. Therefore, if carefully designed, a subsidy-based policy can leverage individuals' private information on their returns and may achieve better welfare than direct policies that neglect this information.

The welfare of subsidy rules has another surprising characteristic besides the advantage of subsidy-based policies compared with direct policies. We show that, under positive selection, the optimal subsidy rule achieves the first-best welfare, which is defined as the highest attainable welfare if the policy-maker can observe the individuals' private information used in the treatment choice decisions. In practice, it is not feasible for a policy-maker to observe private information. However, using subsidies, the policy-maker can simulate the effect of the infeasible policies that target unobserved information. These subsidy rules are the best because when the MTE is decreasing, the subsidy rule can replicate the (infeasible) first-best policy using individuals' self-selection into the treatment.

Next, we identify the optimal policy. Considering the result of welfare representation, the identification problem becomes straightforward if the MTE is known. However, point identification of an entire MTE curve requires the support of a propensity score to be large, which is often impossible. We study two approaches to circumvent this issue. First, under positive selection, we can identify the optimal subsidy provided that we can identify the zero of the MTE curve. Second, we can incorporate other shape restrictions on the MTE and identify a partial ranking among the subsidy rules similar to kasy2016partial.

From a practical perspective, this study addresses the problem of learning the optimal allocation of subsidies from (quasi-) experimental data. Our characterization result simplifies the optimal subsidy problem to the identification of the MTE curve. From this perspective, this study bridges the gap between the problems of policy evaluation and policy design. In particular, the dual role of subsidies emphasized in this study is also relevant to the application. When estimating the effects, subsidies serve as instrumental variables for estimating the treatment effects. When used as policy tools, subsidies can serve as the subject of assignments.

We illustrate the theoretical results by an empirical application. Using the experimental data from groh2016wage, we apply our method to identify the optimal wage subsidies for female college students in Jordan. Although groh2016wage concluded that wage subsidies are ineffective for increasing the long-term labor market participation, our analysis indicates that their conclusion may instead be a consequence of an inefficient allocation of subsidies in the experiment. First, we find that in the experiment, the subsidy is substantially higher than the optimal amount suggested by our method. Thus, individuals with low returns may be negatively selected, leading to a lesser effect of subsidies identified in the experiment. We also find that targeting students' majors can substantially increase efficiency. Specifically, medical school students tend to receive greater benefits from a subsidy, and the amount assigned to them should be different from students of other backgrounds. Therefore, instead of concluding that wage subsidy is ineffective, our result shows that well-designed wage subsidies can be an effective tool for boosting female long-term labor market participation.

This study is presented as follows. The remaining part of this section discusses the literature. Section (ref) introduces the model setup and the subsidy problem. Section (ref) presents the welfare representation results and the optimality conditions for welfare-maximizing subsidy rules. In Section (ref), we compare the optimal welfare of subsidy rules with other types of policies, including direct and infeasible policies. Section (ref) presents a set of results regarding the identifications of optimal subsidies. Section (ref) presents the empirical application. Section (ref) concludes. All proofs are collected in Appendix (ref).

Connection to the literature

This study contributes to a growing literature on personalized treatment rules, including manski2004statistical,dehejia2005program, hirano2009asymptotics,stoye2009minimax,bhattacharya2012inferring,kitagawa2018should,athey2021policy. For a recent review of the subject, see hirano2019statistical. Most studies have focused on the assignment of treatments; however, the case of assigning subsidies as encouragement to treatment is understudied.\footnote{The only exception found is qiu2020optimal, who study the optimal assignment of binary instrumental variables as “encouragements” to treatment take-ups. This study, however, examines the case of a continuous instrument, which is more relevant, for example, when considering monetary subsidies.} Although the existing methods for treatment assignments can be applied to subsidy assignments by adopting an intention-to-treat approach that essentially assigns subsidies based on reduced-form estimates of subsidy effects bhattacharya2012inferring,kitagawa2018should, we argue that a selection-model approach that explicitly models treatment choice as a function of the assigned subsidy can have its advantages. First, our approach allows the realized cost of subsidies to depend on treatment take-ups, which the intention-to-treat approaches cannot as there is no model of treatment take-ups. Second, by explicitly modeling treatment choice, we can incorporate shape restrictions in the selection equation to aid identification and estimation, such as mandating positive effects of the subsidy on treatment take-ups horowitz2017nonparametric. Third, the examined subsidy is continuous, while studies on treatment assignments have mostly focused on the case of a binary treatment.

Although most studies on individualized treatment rules consider policy learning under unconfoundedness, recent studies have examined cases when endogeneity arises for reasons such as noncompliances or omitted variable bias kasy2016partial,cui2020semiparametric,qiu2020optimal,athey2021policy, Byambadalai2021Identification, pu2021estimating. We use instrumental variables for identifications, and our study is no exception. Similar to kasy2016partial, we study the welfare rankings of policies when the treatment effect is only partially identified, although we focus on the assignment of instruments rather than the treatment itself. pu2021estimating introduced IV-optimality, a new notion of optimality, for treatment assignment rules that maximize the worst-case welfare among the identification regions. They also derive a bound on the loss in the welfare of IV-optimal rules relative to the first-best rule, which assigns treatments whenever the effect is positive. This study shows that the optimal subsidy rule outperforms the first-best policy that assigns treatments.

This study analyzes the subsidy rules using the marginal treatment effects (MTE) framework heckman1999local, heckman2001local, heckman2005structural, heckman2007econometric. The MTEs can be identified by the method of local instrumental variables and can be used for predicting the effects of hypothetical policies. Studies have proposed new approaches to its identification and estimation carneiro2009estimating, brinch2017beyond, mogstad2018using, mogstad2020policy, sasaki2021estimation and to apply MTE framework to various research topics, such as unconditional quantile effects martinez2020identification and external validity kowalski2018examine. Among these studies, our study is most closely related to sasaki2020welfare, which also applies the MTE framework to statistical decision rules. However, they focus on the assignment of treatment instead of subsidies. They apply the MTE framework to the method of empirical welfare maximization kitagawa2018should, where policies are assumed to lie in a known policy class restricted to avoid complexity.\footnote{Namely, the policy class has a finite $VC$-dimension.} We do not restrict our candidate policies. Additionally, we emphasize the identification and welfare properties, while sasaki2020welfare emphasized the estimation.

The Model

In this section, we introduce the model's setup, including the optimal subsidy problem faced by a policy-maker. Importantly, the subsidy rule only affects the social welfare through its effect on the behavior response, which can eventually impact the outcome. We begin by introducing the MTE framework.

Data-generating process: the MTE framework

There are two treatment statuses, $1$ and $0$, referred to as treated and untreated, respectively. Let $Y_1$ and $Y_0$ be the potential outcomes under the treatment status $1$ and $0$, respectively. The potential outcomes are related to the observable covariates as

align[align omitted — 93 chars of source]

where $X$ is a vector of the observed covariates that affect the potential outcomes, $\mu_1$ and $\mu_0$ are unknown functions, and $U_1$ and $U_0$ are unobserved random variables. Let $D$ denote the binary variable that indicates the treatment status. Specifically, $D = 1$ if an individual receives the treatment and $D=0$ if an individual does not receive the treatment. The realized outcome is $Y = DY_1 + (1-D)Y_0$.

In our setup, we distinguish two types of instruments. The first type of instrument, denoted by $Z$, is an instrument (or subsidy) randomly assigned in the data but could be manipulated by the policy-maker as policy tools. Specifically, the variable $Z$ has two roles. First, $Z$ is an instrumental variable exogenously set in the data and can facilitate the identification of treatment effects. Second, $Z$ is the subsidy, a policy tool that the policy-maker can use to influence individuals' treatment take-ups.

The second type of instrument, denoted by $W$, is an instrument that only aids in the identification of treatment effects and not subject to the policy-maker's control. Generally, the existence of $W$ can enlarge the identification region of the MTE curve and help identify the optimal subsidy rule. However, to apply our method, one must have a nonmanipulatable instrument $W$.

The treatment take-up is modeled by a latent-index utility model, where the selection into the treatment status depends on the individual characteristics and the instrumental variables. Given $(X,W,Z)$, the treatment take-up $D$ is determined by

align[align omitted — 77 chars of source]

where $U_D$ is the unobserved heterogeneity in the treatment selection process. We can interpret $U_D$ as resistance to treatment take-ups: holding $(X,W,Z)$ fixed, individuals with lower $U_D$ are more likely to select into the treatment. We allow $U_D$ to correlate with $(U_1,U_0)$. That is, the resistance $U_D$ represents an individual's private information regarding the potential outcomes $(Y_1,Y_0)$. The individual uses the private information to aid the treatment choice decision as modeled by Equation ((ref)). In practice, the policy-maker does not observe the private information $U_D$.

The following example illustrates the different variables introduced earlier.

example[Tuition Subsidy] In terms of the tuition subsidy, $Y$ can be considered earnings after graduation, $X$ as individual characteristics such as family background, $D$ as levels of education, $W$ as proximity to colleges, and $Z$ as the tuition for attending public college kane1995labor. While the government has no direct control over students' place of residence, policy-makers may change the tuition subsidy to encourage college enrollment.

The most common assumptions followed in the MTE literature are used in this study.

assumption[Random Assignment] Conditional on $X$, the instrumental variables $(W,Z)$ are independent of the unobserved variables $(U_1,U_0,U_D)$.
assumption[Rank Condition] The propensity score $g(x,w,z)$ is a nontrivial function of $(w,z)$, given any $x$.

As shown by vytlacil2002independence, the MTE model of subsidy characterized by Equations ((ref)) and ((ref)), combined with Assumptions (ref) and (ref), is equivalent to the imbens1994identification assumptions of independence and monotonicity for the local average treatment effect (LATE) interpretation of IV estimands.\footnote{The equivalence result in vytlacil2002independence is derived based on $X = x$. When the monotonicity is global across all values of $x$, chen2021global showed that $g(x,w,z)$ must be additively separable between $x$ and $(w,z)$.}

assumption[Moment Existence] The expectations $E[Y_1]$ and $E[Y_0]$ exist, that is, $E[Y_1] < \infty$ and $E[Y_0]<\infty$.
assumption[Density Existence] The distribution of $U_D$ is absolute continuous with respect to the Lebesgue measure for every $X=x$.

Following Assumption (ref), without loss of generality, we can impose the normalization that $U_D \mid X \sim \text{Unif}[0,1]$, as $g(X,W,Z) \geq U_D$ is equivalent to $F_{U_D \mid X} (g(X,W,Z)) \geq F_{U_D \mid X}(U_D)$, where $F_{U_D \mid X}(u) \equiv \mathbb{P}(U_D \leq u \mid X)$.

The marginal treatment effect (MTE) is defined as

align[align omitted — 113 chars of source]

MTE has two common interpretations: one as an average treatment effect for individuals at different margins and the other as the infinitesimal local average treatment effect (LATE) as it is identified by the local instrumental variable heckman2005structural. MTE corresponds to the change in population outcome resulting from an infinitesimal change in the instrumental variable. The MTE curve can be used as a building block for other conventional treatment effect parameters, such as the average treatment effect. In our analysis, MTE also plays a fundamental role. Subsequently, we can also interpret MTE as the marginal effect of increasing the subsidy. In the first set of our results, we use MTE to characterize social welfare.

By definition, $\text{MTE}(x,u)$ is the mean treatment effect for individuals with $X = x$ at the selection margin $U_D = u$, where a higher $U_D$ implies a lower willingness to select the treatment. Fixing $X = x$, a decreasing MTE curve (along the $u$-dimension) corresponds to the case of positive selection, implying that individuals who benefit more from the treatment are more likely to select it. Positive (or negative) selection can often be motivated by economic theory, econometric specifications, and empirical findings. Typical examples include the choice of education in which individuals with higher returns are more likely to invest in education. While we do not impose the assumption of monotone selection in our main results, we analyze its implications on the characterization, identification, welfare properties, and estimation of optimal policies in the relevant sections throughout this study.

Subsidy rules and counterfactual outcomes

We now describe the policy problem and define the subsidy rules. In our setup, the policy-maker can influence an individual's treatment choice and thereby the realized outcome by manipulating the subsidy $Z$. Formally, a policy $\pi$ is a measurable function

align[align omitted — 88 chars of source]

that maps individual characteristics to the action space $Z^p$.\footnote{Although the instrumental variable $W$ does not affect the potential outcomes, the assignment of subsidy can depend on $W$ as $W$ can affect the treatment take-up. Therefore, the optimal subsidy for type $X = x$ may depend on the value of $W$ as well. It is described further in Section (ref). Furthermore, we restrict our attention to deterministic policies.} The action space $\mathcal{Z}^p$ is the user-specified set of the current subsidy assignments. In some of our results, we assume that the action space is equal to $\mathcal{Z}$, the support of $Z$ in the data; while in others, we allow the action space to be larger than $\mathcal{Z}$. An example of $\mathcal{Z}^p$ would be an interval $\mathcal{Z}^p = [z_l,z_u] \subset \mathbb{R}$, where the range $[z_l,z_u]$ is specified by the policy-maker. We also allow the subsidy to be negative, which represents a tax imposed by the policy-maker. We denote the set of candidate policies as $\Pi$.

Rather than directly setting a mandatory treatment assignment for each individual, a subsidy rule aims at improving the welfare by encouraging individuals to select the treatment with subsidies. Given a policy $\pi$, we assume that the counterfactual treatment choice is

align[align omitted — 89 chars of source]

and the counterfactual outcome is

align[align omitted — 67 chars of source]

A comparison with the case with no policy intervention in Equations ((ref)) and ((ref)) shows the difference that the variable $Z$ is replaced by the subsidy $\pi(X,W)$. Our definition of counterfactual outcomes implicitly assume the following form of policy-invariance: (1) the structural functions $\mu_0,\mu_1,$, and $g$, and (2) the distribution of the economic fundamentals $(X,W, U_1,U_0,U_D)$ would remain the same under the policy intervention $\pi$.

We model the cost of subsidy as follows. Let $c(x,w,z,d)$ be the cost of assigning subsidy $\pi(x,w) = z$ to $(X = x,W = w)$ individuals, who then choose treatment status $D = d$. The counterfactual cost under policy $\pi$ is

align*[align* omitted — 37 chars of source]

Unlike studies that have adopted an intention-to-treat approach kitagawa2018should, our framework allows the realized cost of subsidies to depend on individual's treatment choice $D$. Specifically, the cost can be endogenous to the individual's decision because, unlike the intention-to-treat approaches, the treatment choice $D$ is explicitly modeled. We assume that the cost function $c$ is known to the policy-maker although the realized cost for each individual is ex-ante unknown. The following examples present specific forms of the cost function.

example[Constant Cost] kitagawa2018should studied the optimal eligibility rule for receiving subsidized training in the Job Training Partnership Act program. In their welfare calculation, they imputed the cost of the program as $\$774$ for each eligible individual, regardless of the actual take-up. In the notation of this study, their cost function is effectively $c(x,w,z,d) = c(z)$, which is only a function of the policy assignment. However, as reported in bloom1997benefits, the program take-up varies substantially across different gender and age groups, implying that the realized cost is, in fact, heterogeneous.
example[Voucher Cost] Following Example (ref), a more realistic cost function would be \begin{align*} c(x,w,z,d) = z \cdot d, \end{align*} where $z$ is the amount of subsidy paid by the government. Similar to vouchers, which incur costs only when redeemed, the dependence on the treatment choice $d$ reflects that the subsidy is paid only when the individual attends the program. The heterogeneity of treatment take-up is already embedded in this cost function as both the treatment choice and the propensity score depend on the covariates, as specified in Equation ((ref)).

Following studies on the treatment assignment manski2009identification,kitagawa2018should,athey2021policy, we adopt the additive welfare criterion to evaluate the performance of policies. Given the cost function $c$, the welfare $S(\pi)$ of a policy $\pi$ is defined as

align[align omitted — 95 chars of source]

The additive welfare criterion is flexible such that it can cover various social preferences by suitably transforming the outcome variable. For example, let $\nu$ be a concave function, we can accommodate inequality-averse preference by replacing $Y$ by $\nu(Y)$ atkinson1970measurement.\footnote{By maximizing $\mathbb{E}[Y^\pi]$, the optimal policy necessarily maximizes $V(\mathbb{E}[Y^\pi])$ for any $V$ that is an increasing transformation. The invariance property helps in dealing with cases when welfare cannot be represented by simple averages of individual outcomes. For example, in the insecticide-treated bednet example, where the policy goal is to increase the usage of nets ($Y$), we can incorporate externality by choosing $V$ that maps the coverage rate $\mathbb{E}[Y^\pi]$ to the counterfactual infection rate $V(\mathbb{E}[Y^\pi])$. Such $V$ arguably exists if the externality is approximately determined by the average coverage of nets.} Alternatively, we can interpret the welfare function $\mathbb{E}[\nu(Y^\pi)] - \mathbb{E}[C^\pi] = \mathbb{E}[\nu(Y^\pi) - C^\pi]$ as the average social function in which individuals have quasi-linear preferences.

Given the set of possible subsidy assignments $\mathcal{Z}^p$, the policy class $\Pi$, and the social welfare function $S(\pi)$, a subsidy rule $\pi^*\in\Pi$ is said to be optimal if it attains the social optimum, namely, \[ S(\pi^*) \equiv \sup_{\pi\in\Pi} S(\pi). \] For the ease of expositions, we present our results with cost function $c(x,w,z,d) = z \cdot d$ in the main text. The results for general cost functions can be found in Appendix (ref) and the proofs therein.

remarkThe optimal policy may not be unique. For example, in a simple case where the treatment effect is always zero $(Y_1 = Y_0)$ with zero subsidy cost, any subsidy rule is optimal. Moreover, if any targeting variable is a continuous random one, new optimal rules can be developed by modifying the existing rules on a measure-zero set without changing the implied welfare.\footnote{Therefore, any characterization of optimal policies only apply outside a measure-zero set if at least one of the targeting variable is a continuous random one. }

Optimal Subsidy Rules

This section presents the welfare characterization of subsidy rules using the MTE curve. We use the characterization results to derive optimality conditions of subsidy rules under different scenarios. Simple graphical illustrations are subsequently provided for welfare characterization and optimality conditions.

Necessary conditions for optimality

Our first result shows that the welfare of a policy can be expressed by MTE and the propensity score.

proposition[Characterization of Welfare] Under Assumptions (ref) - (ref), we have \begin{align} \mathbb{E}[Y^\pi] = \mathbb{E}[Y_0] + \mathbb{E} \left[ \int_0^{g(X,W,\pi(X,W))} MTE(X,u) du \right], \end{align} and \begin{align*} \mathbb{E}[C^\pi] &= \mathbb{E}[\pi(X,W)g(X,W,\pi(X,W))] \end{align*} for the cost function $c(x,w,z,d) = z \cdot d$.\footnote{Results for the general cost functions can be found in Appendix (ref).}

The upper bound of the integral of MTE in Equation ((ref)) clarifies the study results. This upper bound is the propensity score induced by the subsidy rule. Mathematically, the subsidy rule affects the welfare by manipulating the integration region of MTE. The optimality conditions in this section are derived by finding the upper bound that maximizes the net area between the MTE curve and the cost function. The welfare properties studied in Section (ref) are also closely related to this upper bound.

As shown in Proposition (ref), the welfare of a subsidy rule has three parts: the baseline outcome for the untreated, the treatment effects on individuals who are induced into treatment status, and the expected costs of subsidies for the treatment takers. The baseline outcome $\mathbb{E}[Y_0]$ is irrelevant for welfare comparison between policies because it is unaffected by the policy.

Here, we examine the methods for finding the optimal policy. The method is to maximize the welfare function point-wise for each combination of $(x,w)$. As implied by the choice Equation ((ref)), if the policy-maker assigns subsidy $\pi(x,w)$ to $(X = x, W = w)$ individuals, the counterfactual take-up rate $u^\pi_{x,w}$ among these individuals would be given by

align*[align* omitted — 139 chars of source]

When the propensity score $g(x,w,z)$ is monotonic in $z$, that is, an increase in subsidy always induces more individuals into the treatment, there is a one-to-one mapping between the amount of subsidy $\pi(x,w)$ and the counterfactual take-up rate $u^\pi(x,w)$. Specifically, the relationship is governed by the propensity score, and the problem can be simplified by changing the variables. We first solve the optimal take-up rate problem

align[align omitted — 195 chars of source]

where $I_{x,w} = \{ g(x,w,z): z \in \mathcal{Z}^p \}$ is the image of $g(x,w,\cdot)$, and $g_{x,w}(z) \equiv g(x,w,z)$. The set $I_{x,w}$ reflects to degree of the policy-maker's influence on treatment take-ups by manipulating the subsidy $Z$, and, $g_{x,w}^{-1}(u)$ is the amount of subsidy needed to induce a take-up rate of $u$. The integration in Equation ((ref)) starts from $0$ as individuals with low $U_D$ are always induced first.

To avoid technical issues, we assume $I_{x,w}$ is a closed set. We also assume that the propensity score $g_{x,w}(z)$ is strictly increasing in $z$ so that $g_{x,w}(z)$ is invertible. The monotonicity of the propensity score can be assumed in the context of subsidies. An increase in the amount of subsidies could always lead to more people participating in the treatment program. The following assumptions ensure that the optimization defined in Equation ((ref)) is well defined and has a solution.

assumption[Continuity] The propensity score $g(x,w,z)$ and the cost function $C(x,w,z)$ are continuous in the subsidy $z$, and $\text{MTE}(x,u)$ is continuous in $u$.
assumption[Invertibility] The propensity score $g(x,w,z)$ is strictly increasing in $z$ for all $(x,w) \in \text{Supp}(X,W)$.

Although both variables $X$ and $W$ are used for targeting, they play different roles in the policy problem because $W$ is excluded from the outcome equation. Unlike $X$, which underlies the heterogeneity in the treatment effect, $W$ only enters the optimization problem through the feasible region $I_{x,w}$ and the propensity score $g(x,w,z)$. In particular, the only value of $W$ as a targeting variable comes from its effect on treatment take-up, which cannot be neglected when the policy provides incentives. In contrast, if the policy were to assign the treatment $D$ instead, targeting variable $W$ is unnecessary as the treatment effect does not depend on $W$, as subsequently discussed.

After finding the optimal take-up rate $u^*_{x,w}$, we can compute the optimal subsidy level $\pi^*(x,w)$ that achieves $u^*_{x,w}$ such that

align[align omitted — 68 chars of source]

By construction, $u^*_{x,w}$ is always achieved by some subsidy level $\pi^*(x,w)\in\mathcal{Z}$ as $u^*_{x,w} \in I_{x,w}$. In fact, $\pi^*(x,w)$ is unique as the propensity score is strictly increasing in the amount of subsidy.

The next proposition summarizes the aforementioned arguments and states the necessary condition for a subsidy rule to be optimal.

proposition[Optimality Condition] Suppose that Assumptions (ref) - (ref) hold, the cost function $c(x,w,z,d) = z \cdot d$, and the action space $\mathcal{Z}^p$ is a closed interval $[z_l,z_u]\subset\mathbb{R}$. If $\pi^*$ is an optimal policy, then $\pi^*$ either satisfies $\Lambda(x,w,\pi^*(x,w)) = 0$, where $\Lambda(x,w,z)$ is defined by \begin{align} \Lambda(x,w,z) \equiv MTE(x,g(x,w,z)) - z - g(x,w,z) \cdot \left[ \frac{\partial}{\partial z}g(x,w,z) \right]^{-1}, \end{align} or satisfies $\pi^*(x,w) \in \{z_l, z_u\}$.

The result for the general cost function $c$ is included in the proof of Proposition (ref) in Appendix (ref). The term $\Lambda(x,w,z)$ is the marginal benefit of subsidy, which is equal to the marginal revenue $\text{MTE}(x,g(x,w,z))$ minus the marginal cost $z + g(x,w,z) \cdot \left[ \frac{\partial}{\partial z}g(x,w,z) \right]^{-1}$, which arises naturally from a monopolist's profit maximization problem. The term $g(x,w,z) \cdot \left[ \frac{\partial}{\partial z}g(x,w,z) \right]^{-1}$ is the elasticity of treatment take-up.

Equation ((ref)) shows that MTE can be interpreted as the average marginal benefit of increasing the amount of subsidy for individuals with $(X = x, W=w)$. Therefore, MTE is not only a treatment effect parameter as commonly understood, but also a policy-relevant parameter in the context of personalized subsidy rule.\footnote{Studies have shown a similar connection between MTE and policy effects. For example, carneiro2010evaluating showed that the marginal policy-relevant treatment effect is a weighted average of MTE. However, in our case, MTE is shown as the marginal effects of subsidies. The distinction appears as we study personalized subsidy rules, whereas studies have considered universal changes in the amount of subsidy.}

Following is a simple example to demonstrate the optimality condition presented in Proposition (ref).

exampleSuppose $X$ is a constant and hence can be omitted. Let the cost $c$ be zero. Let $W$ and $Z$ be supported on $[0,1]$. The treatment response is $g(w,z) = \frac{1}{4}(1+z+w)$. Let the MTE be $\text{MTE}(u) = 4-2u$, which is decreasing. The image of $g(w,\cdot)$ is $I_w = [\frac{1}{4}(1+w),\frac{1}{4}(2+w)]$. Here, the optimal policy is $\pi^*(w) = 1-\frac{3}{5}w$.

Example (ref) shows that although the instrument $W$ is excluded from the outcome equation, $W$ is still valuable for targeting because it affects the selection into treatment.\footnote{See Section (ref) for more discussions on whether to target $W$ for different types of policies.} In the example, the optimal subsidy is decreasing in $w$ for two reasons. First, individuals with high $w$ are ex-ante more likely to select the treatment, therefore requiring less subsidy. Second, as there is positive selection into the treatment ($\text{MTE}$ is decreasing), inducing high-resistance (high $u$) individuals into the treatment status is less consequential.

Sufficient conditions for optimality when MTE is monotone

By definition, $\text{MTE}(x,u)$ is the mean treatment effect for individuals with $X = x$ at the selection margin $U_D = u$, where a higher $U_D$ implies a lower willingness to select the treatment. Fixing $X = x$, a decreasing MTE curve (along the $u$-dimension) corresponds to the case of positive selection, implying that individuals who benefit are more likely to take the treatment.

assumption[Positive Selection] The selection process is said to be positive if MTE$(x,u)$ is weakly decreasing in $u$.
assumption[Negative Selection] The selection process is said to be negative if MTE$(x,u)$ is weakly increasing in $u$.

Empirical evidence supports the monotonicity of MTE. For example, in the context of return to schooling, carneiro2009estimating used the local polynomial regression to obtain a nonparametric estimate of the MTE curve. In their study, figure 3 showed a clear downward-slopping MTE curve. Other empirical evidence includes carneiro2011estimating, cornelissen2018benefits. The monotonicity can also be motivated using both economic theory and econometric specifications. Following are a few such examples.

example[Normal Selection Model] Suppose $Y_1 = X'\beta_1+ U_1$, $Y_0 = X'\beta_0+ U_0$, and $D=\mathbf{1}\{Z'\theta\geq U_D\}$. Further, assume that $(U_1, U_0, U_D)$ is jointly normally distributed and independent of $(X,Z)$, and the variance of $U_D$ is normalized to one. Then $\text{MTE}(x,u) = x'(\beta_1-\beta_0) + (\sigma_{1D}-\sigma_{0D})\Phi^{-1}(u)$, where $\sigma_{1D} = Cov(U_1, U_D), \sigma_{0D}=Cov(U_0,U_D)$, and where $\Phi^{-1}$ is the inverse of the standard normal cumulative function. For extensions to non-normal selection models, see heckman2003simple.
example[Roy Model] In the roy1951some model, the treatment take-up is fully determined by the potential gain, given that $D = \mathbf{1}\{\Delta \geq 0\}$, where $\Delta \equiv Y_1 - Y_0 $. Let $U_D = F_{\Delta}(\Delta)\sim\text{Unif}[0,1]$ be the normalized gain. Then $\text{MTE}(u)=\mathbb{E}[Y_1 - Y_0|U_D=u] = F_\Delta^{-1}(u)$.
example[Generalized Roy Model with Positive Selection] Consider a selection model in which the treatment take-up is partially determined by the potential gain in the form $D = \mathbf{1}\{\phi(X,W,Z,\Delta,V) \geq 0\}$, where $\Delta $ is the individual treatment effect defined in the previous example, and V represents the unobserved heterogeneity. In Appendix (ref), we show that if the function $\phi(X,W,Z,\Delta,V)$ is increasing in $\Delta$, then we can construct a function $g$ and a random variable $U_D\sim\text{Unif[0,1]}$ such that (1) $U_D \perp (W,Z) \mid X$, (2) $D = \mathbf{1}\{g(X,W,Z) \geq U_D\}$, and (3) $\text{MTE} (x,u) = \mathbb{E}[\Delta|X = x, U_D = u]$ is decreasing in $u$.
comment\begin{example} (Monotone Treatment Choice) The monotonicity of the MTE curve can be viewed as a first-difference version of the“monotone treatment selection” (MTS) assumption proposed by manski2000monotone. While the MTS assumption states that the average potential response is higher for those who take the treatment, our monotonicity assumption implies the average gain is higher for those who take the treatment. \end{example}

The following proposition characterizes the optimal policy when the selection is monotone.

commentOur next proposition characterizes the optimal policy when the selection is monotone and when the following assumption regarding the cost function is satisfied. \begin{assumption}[Monotone Cost] The cost function $c(x,w,z,d)$ is weakly increasing in $z$ and $d$. \end{assumption} \begin{assumption}[Concave Propensity]The propensity score $g(x,w,z)$ is weakly concave in $z$. \end{assumption} Assumption (ref) states that the cost must be higher for higher subsidies and that the cost is higher if the individual is induced to the treatment. Our main example of cost $c(x,w,d,z) = z \cdot d$ clearly satisfies Assumption (ref).
proposition[Optimality under Positive Selection] Suppose that Assumptions (ref) - (ref) hold, the action space $\mathcal{Z}^p$ = $[z_l,z_u]$, and $g(x,w,z)$ is weakly concave in $z$. Further assume that the cost function $c(x,w,z,d) = z \cdot d$. Then the optimal subsidy $\pi^*(x,w)$ is given by \begin{equation} \pi^*(x,w) = \begin{cases*} z_{l}, & if $\Lambda(x,w,z) < 0$ for all $z \in [z_l,z_u]$, \\ z^*, & if $\Lambda(x,w,z^*) = 0$ for some $z^*\in [z_l,z_u]$, \\ z_u, & if $\Lambda(x,w,z) > 0 \text{ for all } z \in [z_l,z_u]$, \\ \end{cases*} \end{equation} where $\Lambda(x,w,z)$ is defined in ((ref)).

Proposition (ref) is a complete characterization of the optimal policy as one of the three cases in Equation ((ref)) must hold. When the selection is positive, the marginal return of subsidy decreases because individuals with higher returns are always induced first. Corner solutions arise if the marginal return is always positive or negative. We can also characterize the optimal policy for the case of negative selection, although under the assumption of no cost $c(x,w,z,d) = 0$.

proposition[Optimality under Negative Selection] Suppose that Assumptions (ref) - (ref) and (ref) hold, the action space $\mathcal{Z}^p$ = $[z_l,z_u]$, and $\text{MTE}(x,u)$ is weakly increasing in $u$. Furthermore, assume $c(x,w,z,d) = 0$. Then, the optimal subsidy $\pi^*(x,w)$ is given by \begin{equation} \pi^*(x,w) = \begin{cases} z_l, & if \int^{g(x,w,z_u)}_{g(x,w,z_l)} MTE(x,u) du \leq 0, \\ z_u, & otherwise. \end{cases} \end{equation}

In the case of negative selection, individuals who least benefit from the treatment are always induced first, and the marginal return of subsidy is increasing. Therefore, unless the treatment effect is zero, the optimal subsidy is always a corner solution-it either assigns the highest subsidy under consideration so that individuals with high returns are persuaded to take up the treatment; or it assigns the least amount of subsidy to minimize the potential harm caused by the treatment for low-return individuals.

Graphical Illustration

The results presented in this section are illustrated using graphs. For simplicity, we make the following assumptions: the cost $c = 0$, the potential outcome $Y_0 = 0$, and $X$ and $W$ are constants and hence can be omitted in the discussion. Consequently, the subsidy rule $\pi$ becomes a scalar constant. Under these assumptions, the welfare characterization in Proposition (ref) can be simplified to an integral of MTE from $0$ to $g(\pi)$:

align*[align* omitted — 97 chars of source]

Figure (ref) demonstrates the welfare characterization in Proposition (ref) and the optimality condition in Proposition (ref). Figure (ref) demonstrates the optimal subsidy rule under positive selection.

figure[figure omitted — 3,202 chars of source]
figure[figure omitted — 2,609 chars of source]

Welfare Properties of Subsidy Rules

In most studies, instrumental variables are typically used to identify the treatment effects. In this section, we argue that the instrumental variable has a more fundamental influence on the policy design problem through its function of providing incentives for the treatment take-up. We show that assigning subsidies weakly dominates assigning treatments directly and can achieve the first-best welfare when the MTE is decreasing.

To elaborate, we condition our analysis on $X = x$ throughout this section, implying that $g$ and $\pi$ are only functions of the instrumental variables $W$ and $Z$, and MTE is only a function of $u$. For simplicity, we assume that the action space $\mathcal{Z}^p$ is equal to the support of $Z$.

Subsidies better than mandate

To compare welfare, we introduce direct policies, a new class of policies, that differ from the subsidy rules. The direct policies do not manipulate the subsidy $Z$. Instead, they directly manipulate the treatment take-up. Mathematically, a direct policy is a function $\tau : \text{Supp}(W,Z) \rightarrow \{0,1\}$. For an individual with characteristics $(w,z)$, if $\tau(w,z) = 1$, then the policy-maker makes the treatment mandatory. If $\tau(w,z) = 0$, the individual cannot select the treatment.\footnote{The result in this section can be generalized to allow for the randomization of direct policies. That is, the range of $\tau$ can be convexified to $[0,1]$. For simplicity, we do not consider this convexification in the propositions.} Denote the counterfactual outcome under the direct policy $\tau$ by

align*[align* omitted — 63 chars of source]

We study the following optimal identified welfares under two policy settings:\footnote{In this section, we omit the cost part of the welfare as it is ambiguous to compare the costs of treatment and subsidy without a specific context.}

align*[align* omitted — 303 chars of source]

The subscript “sub” represents “subsidy,” and $S^*_{{\text{sub}}}$ is the identified welfare under subsidy rules. The subscript “dir” represents “direct,” and $S^*_{\text{dir}}$ is the optimal welfare under direct policies. We restrict the comparison of optimal welfare on the set of individuals whose $U_D$ lies in the region $\text{Supp}(g(W,Z))$, on which the treatment effect can be identified. The instrumental variables do not affect the treatment choice of individuals with $U_D$ outside the region $\text{Supp}(g(W,Z))$.\footnote{These individuals are referred to as always-takers or never-takers in the LATE literature.} We exclude these individuals in welfare comparisons because the data are inherently uninformative on the treatment effects and the counterfactual welfare under different policies for these individuals. The next two propositions provide a ranking between the two optimal welfares $S^*_{{\text{sub}}}$ and $S^*_{{\text{dir}}}$.

proposition[Subsidies Better Than Direct Policies] Suppose that Assumptions (ref)-(ref) hold. Then, $S^*_{{\text{sub}}} \geq S^*_{\text{dir}}$. The inequality holds strictly if the supremum in the definition of $S^*_{{\text{sub}}}$ is achieved through a unique policy $\pi^*$, such that $g(W,\pi^*(W))$ lies in the interior of $\text{Supp}(g(W,Z))$ with positive probability.

Proposition (ref) states that the optimal welfare under subsidy rules is always preferred to that under direct policies. Subsidy-based policy weakly dominates the direct policy because the former affects the treatment status through changes in the treatment selection: all else being equal, a small subsidy incentivizes only individuals with low $U_D$ while larger subsidy incentivizes both low- and high-$U_D$ into the treatment status. Intuitively, the subsidy-based policy uses $U_D$ as a targeting variable although $U_D$ is unobservable. This capability to implicitly target with $U_D$ will enhance the welfare because $U_D$ correlates with $(U_1, U_0)$ and thus the potential outcomes.

Proposition (ref) can be mathematically explained. Recall that from the discussion succeeding Proposition (ref), the welfare of a subsidy rule is essentially the net area between the MTE and the cost function from zero to a nontrivial upper bound determined by the subsidy rule. As shown in the proof and in Theorem 1 of sasaki2020welfare, we represent the welfare of a direct policy as an integral of MTE, but the integration region is the unit interval $[0,1]$.\footnote{The cost is not explicitly modeled in sasaki2020welfare. Therefore, the welfare representation is the net area between the MTE and the horizontal axis.} Mathematically, direct policies do not have control over the integration region and therefore not as flexible as subsidy rules. We graphically demonstrate this argument at the end of this section.

We focus on the welfare implications of targeting the instruments $(W,Z)$ in direct policies. The Example (ref) in Section (ref) shows that targeting $W$ is useful in designing subsidy rules. However, when considering direct policies, targeting $W$ or $Z$ does not improve the welfare because $W$ and $Z$ are (conditionally) independent with $(U_1,U_0,U_D)$ and are excluded from the outcome equation. After the treatment probability is assigned, the variation in the instruments is irrelevant to welfare. To formally state this result, we introduce a subclass of direct policies that have constant treatment probability. We call a direct policy $\tau$ a constant policy if $\tau(w,z)$ does not vary with $(w,z)$. The optimal welfare for constant policies is denoted by

align*[align* omitted — 140 chars of source]
proposition[Irrelevance of Instruments in Direct Policies] Suppose that Assumptions (ref)-(ref) hold. Then, $S^*_{\text{dir}} = S^*_{\text{con}}$.

Subsidies achieve first-best welfare

Although subsidy rules have the power to implicitly target on unobserved heterogeneity, targeting with subsidies is not optimal if the policy-maker could observe $U_D$ because the subsidy-based policy only targets individuals in a second-best sense. The targeting is restricted to a specific form that individuals with low $U_D$ have to be in the treatment status whenever the high $U_D$ individuals are. However, we next show that in the specific case of positive selection, the optimal subsidy rule achieves the first-best welfare. That is, the policy-maker cannot further improve the welfare from the optimal subsidy rule even if $U_D$ is observed.

To explain the definition of first-best, we should consider a situation where all the characteristics of an individual are observable, that is, the policy-maker has the power to assign the propensity of treatment based on $(W,Z,U_D)$. We want to keep in mind that targeting on $U_D$ is not feasible in practice. We are only considering such infeasible policies for welfare comparisons. Mathematically, an infeasible policy is a function $\tilde{\tau}: \text{Supp}(W,Z) \times [0,1] \rightarrow \{0,1\}$. The policy-maker uses the information about $(W,Z,U_D)$ to mandate the treatment choice for each individual. Denote the counterfactual outcome under the direct policy $\tau$ by

align*[align* omitted — 96 chars of source]

The first-best welfare is the optimal welfare with respect to the infeasible policies:

align*[align* omitted — 179 chars of source]

where the subscript “fb” represents “first-best.” The following proposition states that subsidy-based policies can achieve the first-best welfare.

proposition[Subsidies can be First-best under Positive Selection] Suppose that Assumptions (ref)-(ref) and (ref) hold. Then, $S^*_{{\text{sub}}} = S^*_{\text{fb}}$.
commentLet $D^{\tilde{\tau}}$ denote the treatment status under the policy $\tilde{\tau}$. Then, $D^{\tilde{\tau}}$ is a Bernoulli random variable with parameter \begin{align*} \mathbb{P}(D^{\tilde{\tau}}=1 \mid W,Z,U_D) = \tilde{\tau}(W,Z,U_D). \end{align*} Similarly, the policy-maker may use the information of $(W,Z,U_D)$ together with an external randomization device to mandate the treatment choice for each individual. We denote $Y^{\tilde{\tau}} = D^{\tilde{\tau}} Y_1 + (1-D^{\tilde{\tau}})Y_0$ as the counterfactual outcome from an infeasible policy ${\tilde{\tau}}$.

The underlying assumption behind Proposition (ref) is that a subsidy-based policy can replicate the treatment assignment of the infeasible policy when the MTE is decreasing. Specifically, the infeasible first-best policy assigns all individuals with $\text{MTE}(u) \geq 0$ to the treatment group. Let $u_x^*$ be a solution to $\text{MTE}(u) = 0$. As individuals with higher MTE are always induced first in the case of positive selection, the subsidy-based policy can achieve the same counterfactual treatment choice if the individuals with $U_D = u_x^*$ are indifferent about the treatment choice under the policy, implying that individuals with $\text{MTE}(u) \geq 0$ are induced to the treatment group.

Graphical Illustration

To demonstrate the welfare comparisons using graphs, we impose the same simplifying assumptions as in Section (ref). We further assume that the support $\text{Supp}(g(W,Z)) = [0,1]$. Under these assumptions, there are two direct policies $\tau=0$ and $\tau=1$. Figure (ref) demonstrates the welfare of the two direct policies, which is either $0$ or the integral of the MTE on the entire unit interval $[0,1]$. Figure (ref) shows that subsidy rules can freely select the integration region. Namely, the MTE is integrated over $[0,g(\pi)]$. Unless $\pi^*$ is a corner solution that belongs to $\{0,1\}$, we can choose a subsidy $\pi$ with $g(\pi) \in (0,1)$ that achieves a strictly higher welfare than the two direct policies.

Here, an infeasible policy $\tilde{\tau}$ is a function from $\text{Supp}(U_D) = [0,1]$ to $\{0,1\}$. To maximize the welfare, the policy-maker would want to mandate anyone with nonnegative MTE$(U_D)$ into treatment and exclude others from the treatment. The optimal infeasible policy $\mathbf{1}\{\text{MTE}(u) \geq 0\}$ achieves the first-best welfare. Figure (ref) shows that the optimal infeasible policy becomes a (feasible) subsidy rule when the MTE decreases.

figure[figure omitted — 2,865 chars of source]
figure[figure omitted — 2,742 chars of source]
commentConsider a (infeasible) direct policy that assigns the propensity of treatment based on both the observed characteristics $(W,Z)$ and the unobserved heterogeneity $U_D$. The corresponding treatment status, denoted as $\widetilde{T}$, is a $\sigma(W,Z,U_D)$-measurable Bernoulli random variable with parameter $\tau_{\widetilde{T}}(W,Z,U_D) = \mathbb{P}(\widetilde{T}=1 \mid W,Z,U_D)$. Let $Y^{\widetilde{T}} = \widetilde{T}Y_1 + (1-\widetilde{T})Y_0$ be the counterfactual outcome from implementing the policy $\widetilde{T}$. We define the optimal welfare $S^*_{\text{inf}}$ with respect to this policy class as \begin{align*} S^*_{inf} & = \sup_{\widetilde{T}: \sigma(W,Z,U_D) -measurable} \mathbb{E} [Y^{\widetilde {T}} \mathbf{1}\{U_D \in Supp(g(W,Z))]. \end{align*} Another class of policies of interest, referred to as “constant policies,” is a collection of direct policies that assign a constant value of propensity for all $(w,z) \in Supp(W,Z)$. Specifically, the class of constant policies consists of direct policies with no personalization. Let $Y^T = TY_1 + (1-T)Y_0$ be the counterfactual outcome from a direct policy $T$. We study the following optimal identified welfare under three policy settings:\footnote{In this section, the cost part of the welfare will be ignored as it is unclear how to compare the cost of assigning treatment and that of assigning subsidies without a specific context.} \begin{align*} S^*_{{sub}} & = \sup_{\pi:\text \sigma(W) \text{-measurable}} \mathbb{E} [Y^\pi \mathbf{1}\{U_D \in \text{Supp}(g(W,Z)) \}], \\ S^*_{\text{dir}} & = \sup_{T:\text\sigma(W,Z) \text{-measurable}} \mathbb{E} [Y^T \mathbf{1}\{U_D \in \text{Supp}(g(W,Z)) \}] ,\\ S^*_{\text{con}} & = \sup_{T:\text\tau_T(w,z) = \tau_T\in[0,1]} \mathbb{E} [Y^T \mathbf{1}\{U_D \in \text{Supp}(g(W,Z))\}]. & \end{align*} The subscript “sub" represents “subsidies,” and $S^*_{{\text{sub}}}$ is the optimal identified welfare when the policy-maker uses subsidy rules. The subscript “dir” represents “direct,” and $S^*_{\text{dir}}$ is the optimal identified welfare when the policy-maker does not manipulate $Z$ but directly assigns treatments based on different values of the instrument.\footnote{Mathematically, this means that $T$ is dependent on $(W,Z)$ through $\tau_T$.} The subscript “con,” denotes “constant,” and $S^*_{\text{con}}$ is the optimal identified welfare when the policy-maker sets the same propensity of treatment for all individuals. By construction, $S^*_{\text{con}} \leq S^*_{\text{dir}}$ because every constant policy is a direct policy. We restrict the comparison of optimal welfare to the set of individuals whose $U_D$ lies in the region $\text{Supp}(g(W,Z))$ on which the treatment effect can be identified. Individuals outside this region are either “always-takers” or “never-takers,” as the exogenous variation in the instrumental variables cannot affect their treatment. Therefore, the data are inherently uninformative about the treatment effects and the counterfactual welfare under different policies for these individuals. Accordingly, we restrict our attention to the identified welfare $\mathbb{E}[Y^\pi\mathbf{1}\{U_D \in \text{Supp}(g(Z,W))\}]$ and $\mathbb{E}[Y^T\mathbf{1}\{U_D \in \text{Supp}(g(Z,W))\}]$. The next two propositions provide a ranking among the optimal welfare under different policy classes.

Identifying the Welfare Ranking

In this section, we consider the identification of MTE and address the issue of identifying the optimal subsidy rule. We first discuss two scenarios in which the optimal policy can be point-identified. We then identify partial ranking among subsidy rules.

By identification, we represent the objects of interest, such as welfare ranking and optimal policy, by the joint distribution of $(Y,D,X,W,Z)$. The analysis is conducted under the full knowledge of observable distributions, namely, the joint distribution of $(Y,D,X,W,Z)$. The welfare ranking $\succsim$ is an ordering on the policy space $\Pi$ such that $\pi\succsim\pi'$, if and only if $S(\pi) \geq S(\pi')$. The identified ranking is deemed partial if for some pairs of policies $(\pi,\pi')$ it is impossible to determine whether $S(\pi) \geq S(\pi')$ or $S(\pi) \leq S(\pi')$, given the observable distribution.

commentLet $\mathscr{P}=\{\pi:(\mathcal{X},\mathcal{W})\rightarrow\mathcal{Z}\}$ denote the set of possible policies. We define a ranking $\succsim$ on $\mathscr{P}$ using the welfare criterion $W(\pi)$, where $\pi \succsim \pi'$ if and only if $S(\pi) \geq S(\pi')$. We define optimal policy as the policy $\pi^* \in \mathscr{Z}$ that maximizes the given welfare criterion, or equivalently, $S(\pi^*) \geq S(\pi), \forall \;\pi\in\mathscr{P}$

Point identification of the optimal subsidy

The welfare representation result indicates that both the welfare ranking and optimal policy are identified if the entire MTE curve and propensity score $g$ are identified. The following corollary presents the results.

corollaryIf for some $(x,w)\in\text{Supp}(X,W)$, $\text{MTE}(x,\cdot)$ is identified on $[0,1]$ and $g(x,w,\cdot)$ is identified on $\mathcal{Z}^p$, then the optimal subsidy $\pi^*(x,w)$ is also identified.

heckman2005structural showed that the MTE can be identified using local instrumental variables (LIV). For completeness, we restate the MTE identification result . Define a function $m$ by

align*[align* omitted — 140 chars of source]

The function $m$ can be identified from the data, given that as the propensity score is identified on its support. The following lemma from heckman2005structural states that $m$ is equal to the MTE on $\textit{Supp}(X,g(X,W,Z))$.

lemma[Identification of MTE] Suppose that Assumptions (ref)-(ref) hold. Further, assume that $g(X,W,Z)$ is a non-degenerate random variable conditional on $X$ and that $0<P(D = 1|X)<1$. Then $\text{MTE}(x,u) = m(x,u), \text{ for all } (x,u) \in \text{Supp}(X,g(X,W,Z))$.

Lemma 1 shows that, to identify the entire MTE curve, the support of the propensity score must cover the unit interval for every $x\in\text{Supp}(X)$. Effectively, a large enough exogenous variation in the propensity score induced by $(W,Z)$ is needed. In literature, this is termed as the large support assumption. Identification of the propensity score $g$ on $\text{Supp}(X,W,Z)$ is straightforward because it is simply the observed probability of take-up, given the covariates and instruments. When $Z^p\not\subset Z$, $g$ on $\text{Supp}(X,W,Z)$ cannot be identified by imposing parametric restriction on the propensity score $g$ such as the probit model.

Point identification of the MTE curve is unnecessary for identifying the optimal policy in two scenarios: First, if the policies under consideration assign subsidies only from the support of $Z$, that is, when $\mathcal{Z}^p \subset \mathcal{Z}$, then the optimal policy is identified. As shown in the next proposition, it is possible to identify the optimal policy without identifying MTE. Define $\mathscr{P}^{id} = \{\pi\in\Pi: \pi(x,w) \in \text{Supp}(Z \mid X=x, W=w) \;\text{for all}\;(x,w)\in \text{Supp}(X,W)\}$ as the set of identiable subsidy rules.

proposition[Identification on the Support] Suppose that Assumptions (ref)-(ref) hold. Then, for any $\pi \in\mathscr{P}^{id}$, \begin{align} \mathbb{E}[Y^{\pi}\mid X,W] = \mathbb{E}[Y\mid X,W,Z=\pi(X,W)], \end{align} and \begin{align} \mathbb{E}[C^\pi\mid X,W] &= \pi(X,W) \cdot \mathbb{E}[D\mid X,W,Z=\pi(X,W)]], \end{align} where $\mathbb{E}[Y|X,W,Z=\pi(X,W)]$ and $\mathbb{E}[D\mid X,W,Z=\pi(X,W)]$ are identified as $\pi(x,w) \in \text{Supp}(Z \mid X=x, W=w)$. Thus, the welfare ranking on $\mathscr{P}^{id}$ is identified.

Proposition (ref) states that the counterfactual welfare can be identified by empirical welfare provided that the subsidy under consideration is observed in the data.

A second scenario in which no point identification of the MTE curve is needed is when individuals positively select into the treatment status. Under the assumption of positive selections, individuals with higher returns are always induced first by the subsidies. Therefore, an amount of subsidy is optimal if the marginal effect of subsidy is zero.

proposition[Identification under Positive Selection] Suppose the assumptions stated in Proposition (ref) hold. If there exists $z^*\in \mathcal{Z}^p$ such that $g(x,w,z^*)$ and $\text{MTE}(x,g(x,w,z^*))$ are identified and that $\Lambda(x,w,z^*) = 0$, then $\pi^*(x,w) = z^*$.

Proposition (ref) states that, if the selection is positive, the optimal amount subsidy can be identified provided that point at which the marginal effect is zero is known. Therefore, no instruments are needed to have large support if it contains the point having a zero marginal effect. Even if the requirement is not met, imposing positive selection still has identification power, as shown in the next subsection.

comment\begin{corollary} Suppose that Assumption (ref), (ref), and (ref) hold. Then, for any $\pi,\pi' \in \mathscr{P}^{id}$, $\pi' \succsim \pi$ if and only if \begin{align} \begin{split} \mathbb{E} \left[ \int_{g(X,W,\pi(X,W))}^{g(X,W,\pi'(X,W))} m(X,u) du \right] &\geq \mathbb{E} \left[ \int_{g(X,W,\pi(X,W))}^{g(X,W,\pi'(X,W))} \pi'(X,W) du \right] \&+ \mathbb{E} \left[ \int_{0}^{g(X,W,\pi(X,W))} \pi'(X,W) - \pi(X,W) du \right]. \end{split} \end{align} \end{corollary}

Partial ranking of subsidy rules

While the results in the previous subsection yielded point identification, the requirements can be restrictive. Therefore, the method for obtaining a partial ranking of subsidy rules by imposing shape restrictions is discussed.

Let $\mathcal{M}^o$ and $\mathcal{G}^o$ be sets of functions that represent, respectively, the functional parameter space of MTE and propensity under possible shape restrictions (e.g., parametric model, monotonicity, and boundedness). The true MTE is assumed to be an element of $\mathcal{M}^o$ and the true propensity is assumed to be an element of $\mathcal{G}^o$. Moreover, the true MTE must coincide with the identifiable function $m$ on the identified region. Therefore, the identified set $\mathcal{M}$ of MTEs under shape restrictions is

align*[align* omitted — 153 chars of source]

The identified set $\mathcal{G}$ of propensities is

align*[align* omitted — 145 chars of source]

Let $\langle \cdot, \cdot \rangle$ be the inner product with respect to the measure that underlies the random vector $(X, U_D)$. Define the dual cone and polar cone of $\mathcal{M}$, respectively, as

align[align omitted — 176 chars of source]

For any subsidy-based policy $\pi$ and propensity $g$, we use $F_{g,\pi}(x,u)$ to denote the conditional cumulative distribution function (CDF) of the propensity score $g(X,W,\pi(X,W))$ given $X = x$, that is, $F_{g,\pi}(x,u) \equiv \mathbb{P}(g(X,W,\pi(X,W)) \leq u\mid X = x)$.

proposition[Identification of Partial Ranking] Suppose that Assumptions (ref)-(ref) hold. For simplicity, assume that the cost $c = 0$. Let $(\pi,\pi')$ be a pair of subsidy rules. If $\{F_{g,\pi'} - F_{g,\pi}: g \in \mathcal{G}\} \subset \mathcal{M}^*$, then $\pi \succsim \pi'$. If $\{F_{g,\pi'} - F_{g,\pi}: g \in \mathcal{G}\} \subset \mathcal{M}^\times$, then $\pi' \succsim \pi$.
comment\begin{align*} \mathcal{M} = & \big\{ \bar{m}: \bar{m} is a function on [0,1] such that \bar{m}(u) = \mathbb{E}[MTE(X,u)] , \\ & where MTE \in \mathcal{M}^o \text{, and } \text{MTE}(x,u) = m(x,u) \text{, for all } (x,u) \in \text{Supp}(X,g(X,W,Z)) \big\} \\ \mathcal{G} = & \left\{ g \in \mathcal{G}^o : g(x,w,z) = \mathbb{E}[D\mid X=x,W=w,Z=z], (x,w,z) \in \text{Supp}((X,Z,W)) \right\}. \end{align*} Let $\langle \cdot, \cdot \rangle$ be the inner product in $L^2([0,1])$ under the Lebesgue measure. Define the dual cone and polar cone of $\mathcal{M}$, respectively, as \begin{align} \mathcal{M}^* & = \{ \ell : \langle l,\bar{m} \rangle \geq 0, \bar{m} \in \mathcal{M} \} \text{, and } \mathcal{M}^\times = -\mathcal{M}^*. \end{align} For any policy $\pi$ and propensity $g$, we use $F_{g,\pi}$ to denote the CDF of the propensity score $g(X,W,\pi(X,W))$. \begin{proposition} [Welfare Ranking] Suppose that Assumption (ref), (ref), and (ref) hold. Additionally, assume that $g \in \mathcal{G}^o$ and $\text{MTE} \in \mathcal{M}^o$. Let $(\pi,\pi')$ be a pair of policies. If $\{F_{g,\pi'} - F_{g,\pi}: g \in \mathcal{G}\} \subset \mathcal{M}^*$, then $\pi \succsim \pi'$. If $\{F_{g,\pi'} - F_{g,\pi}: g \in \mathcal{G}\} \subset \mathcal{M}^\times$, then $\pi' \succsim \pi$. \end{proposition}

Proposition (ref) identifies a partial ranking among the subsidy rules. This result is the subsidy-rules analog of Proposition 1 in kasy2016partial, who studied the identification of a partial ranking among direct policies. The partial identification is achieved by considering the geometry in the Hilbert space containing data-consistent MTE curves. For a pair of subsidy rules, if the difference between the induced CDFs of the propensity score is orthogonal to the set of plausible MTEs, then the data are uninformative about the welfare ranking between the policies. From the decision-theoretic perspective, the identified welfare ranking constitutes an incomplete ordering that admits an expected utility representation dubra2004expected.

Our next proposition states that when there is positive selection, the partial identification result for the policy ranking can be easily interpret.

proposition[Direction of Welfare Improvement] Suppose the assumptions in Proposition (ref) hold, and let $\pi^*$ be an optimal policy. If $\Lambda(x,g(x,w,z)) \geq 0$ (resp. $\leq0$) for some $z \in \mathcal{Z}^p$, then $\pi^*(x,w) \geq z$ (resp. $\leq z$).

The same intuition behind Proposition (ref) applies here: Under positive selection, the marginal benefit of subsidy decreases. Recall that the function $\Lambda(x,w,z)$ only depends on the MTE and the propensity score, and it represents the marginal return of subsidy. Therefore, if the MTE and the propensity score are identified at a specific point $(x,w,z)$, then the optimal subsidy can be bounded from below if the marginal return at that point is positive and can be bounded from above when it is negative.

commentIn practice, there are covariates (observed regressors) $X$ that enter both the treatment choice and outcome equation. We consider our model assumptions ((ref)) and ((ref)) and policies to be defined conditional on $X$. Explaining the conditioning, MTE would become \begin{align*} MTE(x,u) = \mathbb{E} [Y_1 - Y_0 \mid U = u,X = x]. \end{align*} A separability assumption in the MTE \begin{align} MTE(x,u) = \mu(x) + \nu(u) \end{align} is commonly imposed in the literature carneiro2009estimating,carneiro2011estimating,brinch2017beyond,zhou2019marginal. Under such an assumption, MTE is identified on $\text{Supp}(X) \times \text{Supp}(g(Z))$. We define the conditional version of welfare ranking. For simplicity, we assume $\text{Supp}(X)$ is countable. Similar to Definition (ref), \begin{align*} \tilde{Z} \succsim_x \tilde{Z}' \iff \mathbb{E}[\tilde{Y} \mid X = x] \geq \mathbb{E}[\tilde{Y}' \mid X = x]. \end{align*} \begin{proposition} Suppose the separability condition ((ref)) holds for MTE, then the conditional welfare ranking $\succsim_x$ is the same across $x \in \text{Supp}(X)$. Consequently, if $\tilde{Z}$ is an optimal policy with respect to some $x$, then it is the same for any other $x'$. \end{proposition} This result implies that the welfare ranking and optimal policy are independent of $X$. From the identification perspective, if an optimal policy can be identified in $\text{Supp}(g(Z)\mid X = x)$, then the overall optimal policy is identified.

Empirical Application

Our method is illustrated by applying it to the experimental data from the Jordan New Opportunities for Women (Jordan NOW) pilot study. In the experiment, vouchers for wage subsidies are randomly assigned to female college students in their last year of education. These vouchers can be presented to firms while seeking jobs, and if a student with a voucher is employed, the employer can redeem the voucher for up to six months for an amount equal to the minimum wage. The premise of the program is that wage subsidies can help students land their first jobs in which they can acquire experience and skills that can help in their long-term careers. We refer our readers to groh2016wage for more details about the background and their experiment design.

In their study, groh2016wage found that, although wage subsidies substantially increase the employment rate immediately after graduation, their effects on the long-term labor market participation is limited. Specifically, wage subsidies increase the employment rate by about 38 percentage points during the subsidized period. However, the effect vanishes rapidly after the subsidy expires. Seventeen months after the subsidies expire, the effect on employment is less than two percentage points and not statistically significant. groh2016wage concluded that providing wage subsidies is not an effective measure to promote women's long-term labor market participation, at least in the context of Jordan.

In our empirical exercise, we investigate the effectiveness of wage subsidies can be improved through targeting. The welfare function considered is the 30-month earnings ($Y$) after the subsidy period minus the cost of wage subsidy.\footnote{We proxy the 30-month earnings by the monthly earnings reported at the last round, which occurred two years after the voucher had expired.} In the model, $Y$ is the realization of one of the potential outcomes $Y_1$ and $Y_0$, depending on whether the student successfully found a job after graduation during the subsidized period ($D = 1$) or not ($D = 0$). In our policy exercise, we target based on the student's college major. Specifically, we target on whether the student majors in medical assistance ($X$), which includes nursing and pharmacy specializations. We search for the optimal subsidy within the space $\mathcal{Z}^p = [0,900]$, where $z^p = 0$ refers to no subsidy and $z^p = 900$ is the maximal subsidy an individual could receive in the experiment. We use voucher cost in Example (ref).

The optimal amount of subsidy critically depends on two factors: (1) how effectively can wage subsidy encourage and help students to find their first job and (2) the treatment effect of having a first job after graduation on the long-term labor market outcome. These effects are traditionally quantified by imposing joint normality on the error terms $(U_1,U_0,U_D)$ and their independence to $(X,Z)$. Then, the outcome and choice equations are estimated together with either the method of maximum likelihood or the two-step method proposed by heckman1976common. For this illustration, we will proceed accordingly. Although these parametric assumptions could be restrictive, they yield accurate estimates when the sample size is modest. We provide a more flexible method that does not impose normality in Appendix (ref).

Formally, we estimate the following selection model:

align*[align* omitted — 218 chars of source]

where \[ \Sigma =

pmatrix[pmatrix omitted — 169 chars of source]

. \]

table[table omitted — 1,293 chars of source]
figure[figure omitted — 411 chars of source]

Table (ref) presents the estimates from the Heckman two-step method. Similar to the results in groh2016wage, we find that wage subsidies significantly increase the chance of finding a job after graduation. In Figure (ref), we plot the probabilities as functions of subsidies. Notably, at the maximal amount of subsidy, the employment rate almost triples compared with the case with no subsidy. Moreover, the employment rate is higher among students with medical majors at any given level of subsidy. From the policy-maker's perspective, this implies that helping medical students to land their first job is less expensive.

figure[figure omitted — 383 chars of source]

In the normal selection model, heckman2003simple showed that the MTE is \[ \text{MTE}(x,u) = x'(\beta_1 - \beta_0) - (\rho_1\sigma_1 - \rho_0\sigma_0)\Phi^{-1}(u). \] We plot the estimated MTE curves in Figure (ref). Remarkably, first, the MTE curves are downward-sloping, suggesting that individuals positively select into the treatment. Specifically, individuals who are more likely to find a job after graduation tend to benefit more from it for their long-term career prospects. Second, the marginal effect can be negative among individuals with high $U_D$, meaning that wage subsidies can be potentially hard for individuals with low willingness to work.\footnote{A possible explanation for the negative effect is the stigmatization toward voucher users burtless1985targeted.} From the policy-maker's perspective, it is cheaper to help medical students land their first job.

figure[figure omitted — 614 chars of source]

Combining the previous results, we plot the MTE curves and marginal cost of subsidies in Figure (ref). Based on Proposition (ref), we examine the optimal take-up rate at the intersection of the MTE and marginal cost curves. The optimal subsidies substantially differ for the two groups of students. As discussed earlier, as students with medical majors tend to benefit more from subsidies, the optimal take-up rate is higher than that for students from other majors. In fact, by substituting the optimal take-up rate into the inverse of propensity scores, we find that the optimal subsidy for medical students is about JOD $375$, which is approximately one-third of the amount provided in the experiment. However, regarding students from other majors, the estimated optimal subsidy is negative (JOD $-75$). As negative subsidies are excluded in this context, we have a corner solution, which implies that the policy-maker should not provide subsidies for this group of students.

Our results suggest that the subsidy set in the experiments is much higher than their optimal level (in terms of the welfare function we defined). The insignificant long-term effect found in groh2016wage may be a result of excessive amounts of subsidies that may draw individuals with lower returns. The welfare outcome may improve by lowering the amount of subsidy. Moreover, its targeting efficiency can be further enhanced to exploit the heterogeneity in treatment effects and take-ups.

Conclusion

In this study, we examine the problem of allocating subsidies based on individual characteristics. We adopt the MTE framework to analyze the characterization, identification, and welfare properties of subsidy rules. Our results show that subsidy rules generally outperform policies that directly mandate the treatment. In our empirical example, we estimate the optimal wage subsidy using a parametric MTE model. More flexible methods that do not impose parametric assumptions are provided in the appendix. Their theoretical properties are of interest for future studies.

comment\subsection{Implementation with nonparametric shape restrictions} The parametric assumptions, although convenient, could be unrealistic. One remedy is to adopt a semiparametric methodology commonly used in the MTE literature (e.g., lee2009estimating), which keeps the parametric assumption on $Y_1 = g(X,U_1)$ but relaxes the distributional assumption on the unobserved heterogeneity $(U_1,U_D)$. Similarly, we also provide a semiparametric method for policy learning in appendix (ref). However, the semiparametric method still requires parameterization of the outcome equation. In this subsection, we present an alternative approach that imposes parametric assumptions on neither the distribution of the unobserved heterogeneity nor the functional form of the regression model. Instead, we use the monotone treatment response (MTR) assumption manski2000monotone. In the ITN application, it is reasonable to assume the purchase of the net do not drive individuals away from using a net. That is, it is plausible to impose the MTR assumption: \[ Y_1 \geq Y_0. \] Using the welfare representation result (Lemma (ref)), it is obvious that MTR assumption and the monotonicity of the instrument $Z$ together imply that \[ \mathbb{E}[Y|Z=z] \text{ is increasing in z}, \] suggesting that we can estimate $\mathbb{E}[Y|Z]$ using isotonic regression. The policy-maker aims at finding the optimal level of $z$ to maximize \begin{align*} b \cdot \mathbb{E}[Y|X = x, Z = z] - z \cdot g(x,z) \end{align*} for households of type $X = x$. The first term of the expression represents how much the policy-maker values the protection provided by ITNs, given the adoption rate. The second term is the total amount of subsidy paid to the households, which depends on the subsidy per net $z$ and the proportion of buyers $g(x,z)$. Holding $X = x$, we can estimate the welfare function by two isotonic regressions on $\mathbb{E}[Y|Z]$ and $\mathbb{E}[D|Z]$.