EconBase
← Back to paper

Policy Learning under Endogeneity Using Instrumental Variables

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

64,191 characters · 12 sections · 69 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Policy Learning under Endogeneity Using Instrumental Variables

abstractI propose a framework for learning individualized policy rules in observational data settings characterized by endogenous treatment selection and the availability of an instrumental variable. I introduce encouragement rules that manipulate the instrument. By incorporating the marginal treatment effect (MTE) as a policy invariant parameter, I establish the identification of the social welfare criterion for the optimal encouragement rule. Focusing on binary encouragement rules, I propose to estimate the optimal encouragement rule via the Empirical Welfare Maximization (EWM) method and derive the welfare loss convergence rate. I apply my method to advise on the optimal tuition subsidy assignment in Indonesia.

{\em Keywords:} Encouragement rules, selection, marginal treatment effects, empirical welfare maximization, statistical decision rules

Introduction

Policy effects can be heterogeneous, so an important goal for policymakers is to individualize policy interventions to improve social welfare, that is, to assign individuals to different policies based on their observable characteristics. In many cases, policy decisions need to be informed by observational studies or randomized experiments with imperfect compliance, where people endogenously select into treatment. This paper proposes a framework to learn individualized policy interventions in such settings when an instrumental variable (IV) for the treatment is available. Instead of infeasible mandatory treatment rules, I consider a different class of policies that manipulate the instrument, which I refer to as encouragement rules. For example, it is highly costly or impossible to force people to (or not to) attend school. A more realistic scenario is to provide a scholarship or a tuition subsidy, whereas the tuition fee is commonly used as an instrument for school attendance. This paper then studies the identification and estimation of optimal encouragement rules.

To identify optimal encouragement rules, I incorporate the marginal treatment effect (MTE) framework of heckman2005structural to explicitly model the selection into treatment. In this framework, the MTE plays the role of a policy invariant parameter that aids in producing policy counterfactuals. The first main result of this paper is to establish the identification of the social welfare criterion of encouragement rules, which is defined as the average counterfactual outcome, via the identification of the MTE. In this sense, I bridge the literatures on the MTE and on statistical decision rules of policy interventions. More specifically, the social welfare criterion can be represented as a function of the MTE and the propensity score. This social welfare representation has two benefits. First, it helps the policymaker understand how the optimal encouragement rule is driven by heterogeneity in treatment take-up and treatment effects. Second, it suggests a natural way of performing extrapolation. An encouragement rule can induce variation in the propensity score beyond the observed support. This variation represents individuals whose treatment choice is affected not by the observed instrument but by the encouragement rule. Learning about their average outcome requires extrapolation, and point identification can be restored by assuming semiparametric or parametric models for the MTE and the propensity score, depending on the extent of extrapolation required.

To estimate optimal encouragement rules, I apply the social welfare criterion of encouragement rules, identified via the MTE function, to one popular class of statistical decision rules: Empirical Welfare Maximization (EWM) rules hirano2020asymptotic. The EWM approach directly chooses an optimal policy from a constrained class of feasible policies based on sample data. Constraints naturally arise in realistic settings when the policymaker wants to avoid complicated rules or satisfy legal, ethical, or political considerations. The EWM approach has proven to be practically implementable.\footnote{Implementable algorithms include mixed integer linear programming kitagawa2018should and policy trees athey2021policy. EWM rules with mandatory treatment assignment have been implemented in empirical applications using data from the National Job Training Partnership Act (JTPA) Study kitagawa2018should,kitagawa2021equality,mbakop2021model,sasaki2020welfare, the Oregon Health Insurance Experiment (OHIE) sun2021empirical, and the California Greater Avenues for Independence (GAIN) Program athey2021policy, among others.} The second main result of this paper is to establish the convergence rate of the average welfare loss (regret) of the EWM encouragement rule relative to an oracle optimal rule, while allowing for a wide class of nonparametric and parametric estimators for the MTE and the propensity score. To keep the analysis tractable, I focus on settings in which the policymaker allocates individuals to two a priori chosen manipulations of the instrument and constrains the class of feasible allocations.\footnote{In principle, it is possible to generalize the regret analysis to multi-action settings by incorporating different complexity measures of the policy class, such as the entropy integral used in zhou2023offline, but a full analysis is beyond the scope of this paper.}

I further consider two practically relevant extensions. First, when there are multiple instruments, I propose to work with a treatment selection model that allows for unobserved heterogeneity in the marginal rate of substitution across instruments. Second, when there is a budget constraint, I introduce the budget-constrained EWM encouragement rule and analyze its properties in terms of asymptotic optimality and asymptotic feasibility.

I apply the EWM encouragement rule to an empirical dataset from the third wave of the Indonesian Family Life Survey (IFLS). The goal is to provide advice on how upper secondary schooling can be encouraged to maximize average adult wages by manipulating the tuition fee. I find that the optimal policy without budget constraints provides tuition subsidy eligibility to individuals who face relatively high tuition fees and live relatively close to the nearest secondary school. I also provide a partial explanation of why this subpopulation is targeted.

Related Literature: This paper is related to four strands of literature, of which I provide a non-exhaustive overview below.

First, the research question is closely related to the literature on statistical treatment rules in econometrics following the seminal work of manski2004statistical. See hirano2020asymptotic for a recent review. Despite the breadth of the literature, only a few works look into observational data settings when the unconfoundedness assumption does not hold. kasy2016partial and byambadalai2022 focus on cases of partial identification and welfare ranking of policies rather than optimal policy choices. athey2021policy assume homogeneous treatment effects so that the conditional average treatment effect (CATE) on compliers can be extrapolated to those on the entire population. sasaki2020welfare identify the social welfare criterion via the MTE and demonstrate an application to the EWM framework. Nonetheless, these works implicitly assume complete enforcement of treatment rules, whereas I consider more realistic policy tools. In particular, the social welfare representation for treatment rules in sasaki2020welfare can be viewed as a special case of that for binary encouragement rules in this paper when manipulations of the instrument are extremely strong. This echoes the discussion in their Section 3 of the presumption of full compliance under the treatment assignment, which they recognize will be rationalized in extreme circumstances.

The only exception I am aware of is chen2022personalized, who use the MTE framework to study the personalized subsidy rule. Their work is complementary to mine in emphasizing different aspects of policy learning. They focus on the oracle optimal policy without restricting the policy class, whereas I analyze the estimated optimal policy within a restricted policy class. Notably, their closed-form characterization of optimal subsidy rules requires monotonicity of the MTE function with respect to the selection unobservable, which can be restrictive in practice. For example, in the context of the effects of family size on child outcomes, the quantity-quality model of fertility by becker1973interaction is consistent with both positive and negative effects of family size depending on the level of complementarity in parental preferences between quantity and quality of children. Consequently, a monotone MTE function can mask important heterogeneity. Indeed, the MTE estimates in brinch2017beyond show a U shape. In contrast, my framework allows for a flexible form of the MTE function.

Second, in epidemiology and biostatistics, there has been increasing interest in individualized treatment rules. cui2020semiparametric and qiu2020optimal allow for treatment endogeneity. They achieve point identification by leveraging the “no common effect modifier” assumption outlined in wang2018bounded, which largely restricts the heterogeneity of compliance behavior. pu2021estimating introduce the notion of “IV-optimality” to estimate the optimal treatment regime based on partial identification of the CATE. Unlike my framework, these works do not account for imperfect enforcement as a consequence of treatment endogeneity. One exception is qiu2020optimal, who consider individualized encouragement rules that manipulate a binary instrument. My framework nests theirs by allowing the instrument to have richer support.

Third, if one entirely discards information about the treatment and focuses on the relationship between the instrument and the outcome, as in the intention-to-treat analysis, then the optimal manipulation of the instrument can be studied within a policy learning framework with general action spaces. Existing works have covered settings with multivalued actions zhou2023offline,fang2025model and continuous actions kallus2018policy,ai2024data. These works assume unconfoundedness and strong overlap, so that identification of the average counterfactual outcome is not an issue. In contrast, my approach explicitly incorporates information about the treatment and imposes additional structure, namely that the instrument affects the outcome only through the treatment. Under this structure, the average counterfactual outcome depends on the MTE and the propensity score in an interpretable way, which also enables rigorous extrapolation away from the instrument variation observed in the data.

Lastly, the representation of the social welfare criterion in this paper can be viewed as a variation of policy relevant treatment effects (PRTE), adding to the class of policy parameters that can be written as weighted averages of the MTE. Hence, this paper complements the literature on PRTE, including carneiro2010evaluating,carneiro2011estimating, mogstad2018using, and sasaki2021estimation, among many others.

Organization: The rest of the paper is organized as follows. Section (ref) sets up the model, introduces the encouragement rule, and derives a representation of the social welfare criterion via the MTE. It also discusses the identification of the social welfare criterion based on this representation and elaborates on binary encouragement rules. Section (ref) applies the social welfare representation to the EWM method and analyzes the regret properties. Section (ref) discusses extensions incorporating multiple instruments and budget constraints. Section (ref) presents an empirical application. Section (ref) concludes. Proofs and additional results are collected in the appendix.

Encouragement Rules with An Instrument

Setup

I consider the canonical program evaluation problem with a binary treatment $D\in\{0,1\}$ and a scalar, real-valued outcome $Y\in\mathcal{Y}\subset\mathbb{R}$. Outcome production is modeled through the potential outcomes framework rubin1974estimating:

equation*[equation* omitted — 39 chars of source]

where $(Y(0),Y(1))$ are the potential outcomes under no treatment and under treatment.\footnote{By adopting the potential outcomes model, I implicitly follow the conventional practice of imposing the Stable Unit Treatment Value Assumption (SUTVA), namely that there are no spillover or general equilibrium effects.} Let $X\in\mathcal{X}\subset \mathbb{R}^{d_x}$ denote a vector of pretreatment covariates. For instance, in the analysis of returns to schooling, $D$ is an indicator for school enrollment, $Y$ is the log wage, and $X$ includes observable characteristics that affect wages (e.g., parental education, rural/urban residence).

A (non-randomized) treatment rule is defined as a mapping $\pi:\mathcal{X}\to\{0,1\}$. The policymaker’s objective function is the utilitarian (additive) welfare criterion defined by the average counterfactual outcome:

equation*[equation* omitted — 80 chars of source]

Define the conditional average treatment response functions as $\mu_d(x)=E[Y(d)|X=x],d=0,1$. By the law of iterated expectations,

equation*[equation* omitted — 75 chars of source]

Under the unconfoundedness assumption that $D$ is independent of $(Y(0),Y(1))$ conditional on $X$, $\mu_d(x)$ is identified by $E[Y|D=d,X=x]$ for $d=0,1$. However, the unconfoundedness assumption is violated if, for example, people self-select into schooling based on unmeasured benefits and costs driven by ability and motivation, both of which also affect wages. As a result, $\mu_d(x)\neq E[Y|D=d,X=x]$ for $d=0,1$ in general, and thus the social welfare criterion is not identified by the moments of observables. In this case, it is helpful to assume that there exists an instrument (i.e., an excluded variable) $Z\in\mathcal{Z}\subset \mathbb{R}$ that affects the treatment but not the outcome, e.g., the tuition fee. For each $z\in\mathcal{Z}$, denote the potential treatment status if the instrument were set to $z$ by $D(z)$. The observed treatment is given by $D=D(Z)$. I explicitly model the selection into treatment via an additively separable latent index model:

equation[equation omitted — 78 chars of source]

where $\tilde{\nu}$ is an unknown function, and $\tilde{U}$ represents unobservable factors that affect treatment choice. Let “$\perp$” denote (conditional) statistical independence. I adopt the following assumptions from the MTE literature heckman2005structural,mogstad2018using:

enumerate[label=Assumption \arabic*,ref=\arabic*,itemindent=5\parindent,leftmargin=0pt] • (IV Restrictions and Continuous Distribution) \begin{enumerate}[label=(\roman*)] • $\tilde{U}\perp Z|X$. • $E[Y(d)|X,Z,\tilde{U}]=E[Y(d)|X,\tilde{U}]$ and $E[|Y(d)|]<\infty$ for $d\in\{0,1\}$. • $\tilde{U}$ is continuously distributed conditional on $X$. \end{enumerate}

Assumptions (ref)(i) and (ii) impose exogeneity and an exclusion restriction on $Z$ but allow for arbitrary dependence between $(Y(0),Y(1))$ and $\tilde{U}$, even conditional on $X$. vytlacil2002independence shows that, under Assumption (ref)(i), the existence of an additively separable selection equation as in ((ref)) is equivalent to the monotonicity assumption used for the local average treatment effects (LATE) model of imbens1994identification. The LATE monotonicity assumption restricts choice behavior in the sense that, conditional on $X$, an exogenous shift in $Z$ either weakly encourages or discourages every individual to choose $D=1$. Nonetheless, I maintain the selection equation ((ref)) because it allows me to express the average counterfactual outcome as a function of identifiable and interpretable objects. Under Assumptions (ref)(i) and (iii), one can reparameterize the model as

equation[equation omitted — 124 chars of source]

where $U\equiv F_{\tilde{U}|X}(\tilde{U}|X)$ and $\nu(x,z)\equiv F_{\tilde{U}|X}(\tilde{\nu}(x,z)|x)$. As a consequence,

equation*[equation* omitted — 79 chars of source]

where $p(x,z)$ is the propensity score.

Endogenous treatment selection also challenges the plausibility of fully mandating treatment assignment. I instead consider a different class of policies that manipulate the instrument, which I refer to as encouragement rules. Formally, an encouragement rule is a mapping $\boldsymbol{\alpha}:\mathcal{X}\times\mathcal{Z}\to\mathbb{R}$ that manipulates the instrument for an individual with $(X,Z)=(x,z)$ from the initial value $z$ to a new level $\boldsymbol{\alpha}(x,z)$. For example, when $Z$ is the tuition fee, $\boldsymbol{\alpha}(x,z)=(z-\alpha(x))\cdot1\{z\geq\alpha(x)\}$ with $\alpha:\mathcal{X}\to\mathbb{R}_+$ describes a tuition subsidy rule that subsidizes an individual with $X=x$ up to $\alpha(x)$. The representation of the social welfare criterion in Section (ref) covers the most general setting without restrictions on $\boldsymbol{\alpha}$. When I study the regret bounds for statistical decision rules in Section (ref), I take a stand on the complexity of feasible encouragement rules. In particular, I focus on a case in which the policymaker allocates individuals to two a priori chosen manipulations of the instrument. The binary formulation of encouragement rules nests treatment rules kitagawa2018should,sasaki2020welfare as a special case.

Representation of the Social Welfare Criterion via the MTE

The outcome that would be observed under encouragement rule $\boldsymbol{\alpha}$ is

equation*[equation* omitted — 132 chars of source]

I define the social welfare criterion as the average counterfactual outcome:

equation*[equation* omitted — 84 chars of source]

Theorem (ref) shows that $W(\boldsymbol{\alpha})$ can be expressed as a function of the average treatment effect conditional on observable characteristics $X$ and the selection unobservable $U$, which is the definition of the MTE:

equation*[equation* omitted — 72 chars of source]

A proof is provided in (ref). The concept of MTE was introduced by bjorklund1987estimation and extended by heckman2005structural,heckman2007econometric. In contrast to the intention-to-treat approach, the representation in Theorem (ref) has a straightforward interpretation: among the individuals with $(X,Z)=(x,z)$, those for whom the value $u$ of $U$ is between $p(x,\boldsymbol{\alpha}(x,z))$ and $p(x,z)$ get either encouraged or discouraged to take up the treatment, and their contribution to the welfare contrast $W(\boldsymbol{\alpha})-E[Y]$ is $\operatorname{MTE}(x,u)$ if encouraged and $-\operatorname{MTE}(x,u)$ if discouraged.

theoremUnder Assumption (ref), the social welfare criterion for a given encouragement rule $\boldsymbol{\alpha}$ is given by \begin{equation*} W(\boldsymbol{\alpha})=E[Y]+E\Big[\int_0^1 \operatorname{MTE}(X,u)\cdot(1\{p(X,\boldsymbol{\alpha}(X,Z))\geq u\}-1\{p(X,Z)\geq u\})\,du\Big]. \end{equation*}

A similar representation result appears in chen2022personalized, where they exclude $Z$ from the targeting variables. This exclusion narrows their focus to policies that induce a degenerate distribution of $Z$ conditional on $X=x$. In contrast, my framework also accommodates policies that shift the conditional distribution of $Z$ through deterministic transformations, e.g., $\boldsymbol{\alpha}(x,z)=z+\alpha(x)$ with $\alpha:\mathcal{X}\to\mathbb{R}$.

It is worth noting that $W(\boldsymbol{\alpha})$ has a natural connection to the concept of policy relevant treatment effects (PRTE) heckman2005structural. For a general class of policies that affect the propensity score, the PRTE is defined as the mean effect of going from a baseline policy to an alternative policy per net person shifted (assuming that $E[D|\text{alternative policy}]-E[D|\text{baseline policy}]\neq0$):

equation*[equation* omitted — 147 chars of source]

Corollary (ref) gives an alternative representation of $W(\boldsymbol{\alpha})$ in terms of a suitably defined PRTE.

corollaryUnder Assumption (ref), when $E[p(X,\boldsymbol{\alpha}(X,Z))]-E[p(X,Z)]\neq0$, the social welfare criterion for a given encouragement rule $\boldsymbol{\alpha}$ is given by \begin{equation*} W(\boldsymbol{\alpha})=E[Y]+(E[p(X,\boldsymbol{\alpha}(X,Z))]-E[p(X,Z)])\cdot\operatorname{PRTE}(\boldsymbol{\alpha}), \end{equation*} where \begin{equation*} \operatorname{PRTE}(\boldsymbol{\alpha})=E\Big[\int_0^1 \operatorname{MTE}(X,u)\cdot \omega(X,Z,u;\boldsymbol{\alpha})\,du\Big] \end{equation*} with the weight defined as \begin{equation*} \omega(x,z,u;\boldsymbol{\alpha})\equiv\frac{1\{p(x,\boldsymbol{\alpha}(x,z))\geq u\}-1\{p(x,z)\geq u\}}{E[p(X,\boldsymbol{\alpha}(X,Z))]-E[p(X,Z)]}. \end{equation*}

Corollary (ref) unfolds the two forces driving the optimal policy based on $W(\boldsymbol{\alpha})$: the average change in treatment take-up, $E[p(X,\boldsymbol{\alpha}(X,Z))]-E[p(X,Z)]$, and the average treatment effect among those induced to switch treatment status, $\operatorname{PRTE}(\boldsymbol{\alpha})$, when going from the status quo to encouragement rule $\boldsymbol{\alpha}$. Moreover, $\operatorname{PRTE}(\boldsymbol{\alpha})$ can be expressed as a weighted average of the MTE with weights determined by both observed and unobserved heterogeneity in treatment take-up.

Identification of the Social Welfare Criterion

Theorem (ref) implies that the point-identification of $W(\boldsymbol{\alpha})$ is guaranteed by the point-identification of the propensity score and the MTE over necessary domains. I formalize this insight in the following assumption.

enumerate[label=Assumption \arabic*,ref=\arabic*,itemindent=5\parindent,leftmargin=0pt] \setcounter{enumi}{1} • (Point-Identification of $W(\boldsymbol{\alpha})$) \begin{enumerate}[label=(\roman*)] • $p(x,z)$ is point-identified over $\operatorname{Supp}(X,\boldsymbol{\alpha}(X,Z))$. • For every $x\in\mathcal{X}$, $\operatorname{MTE}(x,\cdot)$ is point-identified over $[\min\mathcal{P}_{\boldsymbol{\alpha}}(x),\max\mathcal{P}_{\boldsymbol{\alpha}}(x)]$, where $\mathcal{P}_{\boldsymbol{\alpha}}(x)$ denotes the support of $p(X,\boldsymbol{\alpha}(X,Z))$ conditional on $X=x$. \end{enumerate}

The method of local instrumental variables (LIV) heckman1999local,heckman2001local gives

equation[equation omitted — 100 chars of source]

provided that $u\mapsto E[Y|X=x,p(X,Z)=u]$ is continuously differentiable for almost every $x$. Therefore, Assumption (ref)(ii) is satisfied via the point-identification of the derivative of $E[Y|X=x,p(X,Z)=u]$ with respect to $u$. As a result, $W(\boldsymbol{\alpha})$ is point-identified as

equation*[equation* omitted — 107 chars of source]

where $\mu_Y(x,u)\equiv E[Y|X=x,p(X,Z)=u]$.

I give three examples of sufficient conditions for Assumption (ref). Verification of Assumption (ref) in these examples is relegated to (ref). Typically, additional structural restrictions are required to compensate for the relaxation of the support condition.

example[Nonparametric Identification] Assume that (E1-1) $\operatorname{Supp}(X,\boldsymbol{\alpha}(X,Z))\subset\operatorname{Supp}(X,\allowbreak Z)$; (E1-2) the conditional distribution of $p(X,Z)$ given $X$ is absolutely continuous with respect to the Lebesgue measure. Then, Assumption (ref) is satisfied.
example[Semiparametric Identification] Assume that (E2-1) $\operatorname{Supp}(\boldsymbol{\alpha}(X,Z))\subset\mathcal{Z}$; (E2-2) the distribution of $p(X,Z)$ is absolutely continuous with respect to the Lebesgue measure; (E2-3) the propensity score is modeled as $p(x,z)=x^\top\gamma+\theta(z)$, where $\gamma$ is an unknown parameter and $\theta$ is an unknown function; (E2-4) the potential outcomes are modeled as $Y(0)=X^\top\beta_0+V_0$ and $Y(1)=X^\top\beta_1+V_1$, where $\beta_1$ and $\beta_0$ are unknown parameters, and $E[V_d|X,Z,U]=E[V_d|U]$ for $d=0,1$;\footnote{The partially linear form for potential outcomes in (E2-4) are commonly assumed in applied work estimating the MTE; see, e.g., carneiro2009estimating, carneiro2010evaluating,carneiro2011estimating. These works also invoke full independence $(U,V_0,V_1)\perp(Z,X)$, which is stronger than the conditional mean independence of $(V_0,V_1)$ from $(Z,X)$ in (E2-4).} (E2-5) $E[(X-E[X|Z])(X-E[X|Z])^\top]$ and $E[(X-E[X|p(X,Z)])(X-E[X|p(X,Z)])^\top]$ are positive definite. Then, Assumption (ref) is satisfied. When (E2-1) is violated, Assumption (ref) can still be satisfied by imposing a parametric model for the propensity score and a semiparametric partially linear model for the MTE, which is what I implement in the empirical application.
example[Parametric Identification] Assume that (E3-1) the propensity score is modeled as $p(x,z)=b(x,z)^\top\gamma$, where $b$ is a known vector function and $\gamma$ is an unknown parameter;\footnote{Alternatively, one may adopt a logit or probit model to respect the $[0,1]$ boundary.} (E3-2) the conditional mean of $Y$ given $(X,p(X,Z))$ is modeled as $E[Y|X=x,p(X,Z)=u]=ux^\top\beta_1+(1-u)x^\top\beta_0+\sum_{j=2}^J \eta_ju^j$, where $\beta_1$, $\beta_0$, and $\eta_2,\dots,\eta_J$ are unknown parameters; (E3-3) there is no multicollinearity in $b(X,Z)$ elements nor in $(p(X,Z)X^\top,\allowbreak(1-p(X,Z))X^\top,p(X,Z)^2,\allowbreak\dots,p(X,Z)^J)$.\footnote{Polynomial MTE models are often used in empirical studies; see, e.g., brinch2017beyond, cornelissen2018benefits.} Then, Assumption (ref) is satisfied.

Binary Encouragement Rules

When studying the regret bounds for statistical decision rules in Section (ref), I focus on settings in which the policymaker allocates individuals to two a priori chosen manipulations of the instrument. Formally, given two functions $\alpha_0,\alpha_1:\mathcal{X}\times\mathcal{Z}\to\mathbb{R}$, a binary encouragement rule, indexed by a mapping $\pi:\mathcal{X}\times\mathcal{Z}\to\{0,1\}$, manipulates the instrument for an individual with $(X,Z)=(x,z)$ to

equation*[equation* omitted — 104 chars of source]

Corollary (ref) specializes Theorem (ref) to the binary setting.

corollaryUnder Assumption (ref), the social welfare criterion for a given binary encouragement rule $\pi$ is given by \begin{eqnarray*} W(\boldsymbol{\alpha}^\pi)&=&E[Y(D(\alpha_0(X,Z)))]+E\Big[\pi(X,Z)\\ &&\cdot\int_0^1 \operatorname{MTE}(X,u)\cdot(1\{p(X,\alpha_1(X,Z))\geq u\}-1\{p(X,\alpha_0(X,Z))\geq u\})\,du\Big]. \end{eqnarray*}

Finally, I demonstrate that the binary formulation of encouragement rules nests as a special case treatment rules that directly assign individuals to a certain treatment status. Suppose that $\alpha_0$ and $\alpha_1$ satisfy

equation*[equation* omitted — 99 chars of source]

so that $\alpha_1$ (resp. $\alpha_0$) creates perfectly strong incentives (resp. disincentives) to be in the treated state ($D=1$) across heterogeneous covariate values. In this case, encouragement rules are effectively treatment rules: $D(\boldsymbol{\alpha}^\pi(X,Z))=\pi(X,Z)$.\footnote{In this case, $Z$ is redundant as a targeting variable.} Therefore, Corollary (ref) provides a representation of the social welfare criterion for treatment rules via the MTE function:

equation*[equation* omitted — 112 chars of source]

This representation coincides with Theorem 1 of sasaki2020welfare. However, such powerful manipulations are hard to justify in practice. For example, consider a selection of the form $D=1\{Z\geq \tilde{U}\}$ with $\tilde{U}$ having full support on $\mathbb{R}$ and a manipulation of the form $\alpha_d(x,z)=z+a_d$ for $d=0,1$. Then, one needs to set $a_1=\infty$ and $a_0=-\infty$ to induce full compliance. Indeed, in Section 3 of sasaki2020welfare, they recognize that the presumption of full compliance under the treatment assignment will be rationalized in extreme circumstances such as strong legal power or a large amount of resources held by the policymaker.

Applications to EWM and Regret Properties

In this section, I restrict attention to binary encouragement rules described in Section (ref). I apply the social welfare criterion, identified via the MTE function, to the EWM framework and investigate the theoretical properties of the resulting statistical decision rules.

Some extra notations are needed to facilitate the discussion. Suppose the policymaker observes a random sample $A_i=(Y_i,D_i,X_i,Z_i)$ of size $n$. Let $E_n$ denote the sample average operator, i.e., $E_n f=\frac{1}{n}\sum_{i=1}^n f(A_i)$ for any measurable function $f$. Let $a\vee b=\max\{a,b\}$.

For notational simplicity, denote an encouragement rule and its social welfare by $\pi$ and $W(\pi)$ in place of $\boldsymbol{\alpha}^\pi$ and $W(\boldsymbol{\alpha}^\pi)$, respectively. Let $\Pi$ denote the class of encouragement rules the policymaker can choose from. In view of Corollary (ref) and ((ref)), $W(\pi)$ is point-identified under Assumption (ref) as

equation*[equation* omitted — 127 chars of source]

Define the welfare contrast relative to the baseline policy that allocates everyone to $\alpha_0$ as $\bar{W}(\pi)=W(\pi)-E[Y(D(\alpha_0(X,Z)))]$. The optimal encouragement rule is given by $\pi^*\in\operatorname*{arg\,max}_{\pi\in\Pi}\bar{W}(\pi)$ if the distribution of $(X,Z)$ and the mappings $(x,z)\mapsto p(x,z)$ and $(x,u)\mapsto \mu_Y(x,u)$ are known. However, these quantities are unknown in practice. Given an estimator $\hat{p}(x,z)$ for $p(x,z)$ and an estimator $\hat{\mu}_Y(x,u)$ for $\mu_Y(x,u)$, I construct the empirical welfare criterion $\hat{W}_n(\pi)$ by plugging in these estimators:

equation*[equation* omitted — 142 chars of source]

Then, I define the feasible EWM encouragement rule as

equation[equation omitted — 116 chars of source]

In line with the literature on statistical treatment rules manski2004statistical,kitagawa2018should, I evaluate the performance of an encouragement rule $\pi$ by its regret defined as the welfare loss relative to the highest attainable welfare within class $\Pi$:

equation*[equation* omitted — 56 chars of source]

To analyze the regret of $\hat{\pi}_{\mathrm{FEWM}}$, I impose the following assumptions.

enumerate[label=Assumption \arabic*,ref=\arabic*,itemindent=5\parindent,leftmargin=0pt] \setcounter{enumi}{2} • (Boundedness and Vapnik-Chervonenkis (VC)-Class) \begin{enumerate}[label=(\roman*)] • There exists $\bar{M}<\infty$ such that $\sup_{(u,x)\in[0,1]\times\mathcal{X}}|\operatorname{MTE}(x,u)|\leq\bar{M}$. • $\Pi$ has a finite VC-dimension. \end{enumerate}

Assumption (ref)(i) requires the MTE to be uniformly bounded in $u$ and $x$. Assumption (ref)(ii) controls the complexity of the class $\Pi$ of candidate encouragement rules in terms of VC-dimension. Interested readers can refer to van1996weak for the definition and textbook treatment of VC-dimension. I now give two examples of $\Pi$ that satisfy Assumption (ref)(ii).

example[Linear Eligibility Score (LES)] Let $v\in\mathbb{R}^{d_v}$ be a subvector of $(x,z)$. Consider the class of binary decision rules based on linear eligibility scores: \begin{equation*} \Pi_{\mathrm{LES}}=\{\pi:\pi(x,z)=1\{\lambda_0+\lambda^\top v\geq0\},(\lambda_0,\lambda^\top)\in\mathbb{R}^{d_v+1}\}. \end{equation*} For example, individuals are assigned scholarships if a linear function of their tuition fee and distance to school exceeds some threshold. The EWM method searches over all possible linear coefficients. The VC-dimension of $\Pi_{\mathrm{LES}}$ is $d_v+1$.
example[Threshold Allocations (TA)] Consider the class of binary decision rules based on threshold allocations: \begin{equation*} \Pi_{\mathrm{TA}}=\{\pi:\pi(x,z)=1\{\sigma_k v_k\leq\bar{v}_k for k\in\{1,\dots,d_v\}\},\bar{v}\in\mathbb{R}^{d_v},\sigma\in\{-1,1\}^{d_v}\}. \end{equation*} For example, individuals are assigned scholarships if their tuition fee and distance to school are above or below some thresholds. The EWM method searches over all possible thresholds and directions. The VC-dimension of $\Pi_{\mathrm{TA}}$ is $d_v$.

I also propose the following assumption about the unknown components that show up in the social welfare criterion: the propensity score $p(x,z)$ and the observed conditional average outcome $\mu_Y(x,u)$.

enumerate[label=Assumption \arabic*,ref=\arabic*,itemindent=5\parindent,leftmargin=0pt] \setcounter{enumi}{3} • (Estimation of the Propensity Score and the Observed Conditional Average Outcome) \begin{enumerate}[label=(\roman*)] • There exists a sequence $\psi_n\to\infty$ such that for each $d\in\{0,1\}$, \begin{equation*} E[|\hat{p}(X,\alpha_d(X,Z))-p(X,\alpha_d(X,Z))|]=O(\psi_n^{-1}). \end{equation*} • There exists a sequence $\phi_n\to\infty$ such that \begin{equation*} E\Big[\sup_{u\in[0,1]}|\hat{\mu}_Y(X,u)-\mu_Y(X,u)|\Big]=O(\phi_n^{-1}). \end{equation*} \end{enumerate}

Assumption (ref)(i) concerns the convergence rate in expectation of the estimation error for $p(x,z)$. When $p(x,z)$ is estimated nonparametrically as in Example (ref), a sufficient condition for Assumption (ref)(i) is $E[\sup_{(x,z)\in\operatorname{Supp}(X,Z)}|\hat{p}(x,z)-p(x,z)|]=O(\psi_n^{-1})$. In (ref), I derive the sup-norm convergence rate in expectation for local polynomial estimators and series estimators built on exponential tail bounds. The rate can be faster if a semiparametric or parametric estimator is used under additional assumptions as in Example (ref) or (ref). Assumption (ref)(ii) concerns the convergence rate in expectation of the estimation error for $\mu_Y(x,u)$. Since $U$ is not observed, I take the supremum over the unit interval. Usually, $\mu_Y(x,u)$ is estimated using the estimated propensity score as a generated regressor. When the regression model is parametric, I provide sufficient conditions for Assumption (ref)(ii) to hold with $\phi_n=\psi_n$ in (ref). When the regression model is nonparametric, the sup-norm convergence rate in probability is established in mammen2012nonparametric. However, the sup-norm convergence rate in expectation remains unknown. I leave it for future work.

theoremSuppose that Assumptions (ref)-(ref) hold. Then, \begin{equation*} E[R(\hat{\pi}_{\mathrm{FEWM}})]=O(\psi_n^{-1}\vee\phi_n^{-1}\vee n^{-1/2}). \end{equation*}

Theorem (ref) derives a convergence rate upper bound for the average regret of $\hat{\pi}_{\mathrm{FEWM}}$. A proof is provided in (ref). In general, the convergence rate upper bound is determined by $\psi_n^{-1}\vee\phi_n^{-1}$.

remarkThere are two special cases where the $n^{-1/2}$ rate can be achieved.\footnote{While the convergence rate lower bound in the current context is not known, it is natural to conjecture that it is $O(n^{-1/2})$. I leave the formal analysis for future work.} One is to assume parametric forms for $p(x,z)$ and $\mu_Y(x,u)$ so that $\phi_n=\psi_n=n^{1/2}$. The other is to pursue a doubly robust approach in the spirit of athey2021policy. The idea is to use an alternative social welfare criterion based on a doubly robust score, which is Neyman-orthogonal with respect to $p(x,z)$ and $\mu_Y(x,u)$. The details are given in (ref). An extra cost to pay for implementing the doubly robust score is the estimation of the joint density of $(X,Z)$, which can be challenging if the dimension of $X$ is large.

Extensions

I consider two empirically relevant extensions to the baseline setup in Section (ref). In Section (ref), I allow for the presence of other instruments in addition to the one that can be manipulated. In Section (ref), I incorporate budget constraints. As a further extension, I consider encouragement rules with a binary instrument in (ref).

Multiple Instruments

In practice, the policymaker can observe multiple instruments, but only one of them can be used as the tool for policy intervention. For example, tuition subsidies and proximity to upper secondary schools are two instruments for enrollment in upper secondary school, but only the former can serve as an encouragement. More generally, I allow $Z$ to be $L$-dimensional. Let $Z_1\in\mathcal{Z}_1\subset\mathbb{R}$ be the instrument that can be intervened upon, and let $Z_{-1}$ collect all other $(L-1)$ components. An encouragement rule is a mapping $\boldsymbol{\alpha}:\mathcal{X}\times\mathcal{Z}\to\mathbb{R}$ that determines the manipulated level of $Z_1$ while leaving $Z_{-1}$ unchanged.

For $z_1\in\mathcal{Z}_1$, I construct a selection equation for the potential treatment status if $Z_1$ were set to $z_1$ while $Z_{-1}$ remained at its observed realization as

equation[equation omitted — 151 chars of source]

where $U_1$ can be interpreted as a latent proneness to take the treatment, which is measured against the incentive (or disincentive) created by the manipulated instrument.\footnote{Since only one instrument is manipulated, I focus on the “marginal” selection behavior induced by this instrument conditional on the other instruments. If one is interested in policies that simultaneously manipulate multiple instruments, then a treatment selection model with multidimensional unobserved heterogeneity may be needed; see, e.g., ura2024policy.} By ((ref)), I only impose restrictions along one margin of selection and thus are agnostic about unobserved heterogeneity in the marginal rate of substitution across instruments.\footnote{mogstad2021causal use a random utility model to demonstrate that in the presence of multiple instruments, ((ref)) implies homogeneity in the marginal rate of substitution. In contrast, ((ref)) does not impose such implicit homogeneity.} chen2022personalized adhere to ((ref)) when they deal with multiple instruments, thereby presenting the same social welfare representation as in the single-instrument case (i.e., Theorem (ref)).

The outcome that would be observed under encouragement rule $\boldsymbol{\alpha}$ is

equation*[equation* omitted — 152 chars of source]

Define the social welfare criterion as $W(\boldsymbol{\alpha})=E[Y(D(\boldsymbol{\alpha}(X,Z),Z_{-1}))]$. Since ((ref)) only imposes that $U_1$ is independent of $Z_1$ given $(X,Z_{-1})$, I accordingly replace Assumption (ref)(ii) with an exclusion restriction that only requires potential outcomes to be mean independent of $Z_1$ given $(X,Z_{-1})$.

enumerate[label=Assumption \arabic*,ref=\arabic*,itemindent=5\parindent,leftmargin=0pt] \setcounter{enumi}{4} • (Instrument-Specific Exclusion Restriction) $E[Y(d)|X,Z,U_1]\allowbreak=E[Y(d)|X,\allowbreak Z_{-1},U_1]$ and $E[|Y(d)|]<\infty$ for $d\in\{0,1\}$.

It turns out that $W(\boldsymbol{\alpha})$ can be expressed as a function of the instrument-specific MTE defined as

equation*[equation* omitted — 101 chars of source]

The expression is given in Corollary (ref). The analysis in Section (ref) then applies. Heuristically, $\operatorname{MTE}_1$ is equivalent to the MTE function using the manipulated instrument and conditioning on the other instruments as covariates.

corollaryUnder ((ref)) and Assumption (ref), the social welfare criterion for a given encouragement rule $\boldsymbol{\alpha}$ is given by \begin{eqnarray*} W(\boldsymbol{\alpha})&=&E[Y]+E\Big[\int_0^1\operatorname{MTE}_1(X,Z_{-1},u_1)\\ &&\cdot(1\{p(X,\boldsymbol{\alpha}(X,Z),Z_{-1})\geq u_1\}-1\{p(X,Z)\geq u_1\})\,du_1\Big]. \end{eqnarray*}

Budget Constraints

Manipulating the instrument can be costly, especially when the instrument is a monetary variable such as price. In practice, the policymaker often faces budget constraints and wants to prioritize encouragement for the individuals who will benefit the most. Incorporating budget constraints is of particular interest when the treatment effect is intrinsically positive. For example, dupas2014short documents an experiment in Kenya that randomly assigned subsidized prices for a new health product. The treatment and outcome were indicators for the product's purchase and usage, respectively. The product was not available outside the experiment, so the potential outcome if not treated is identically equal to zero. Hence, the first-best decision rule was to assign the treatment, or an encouragement that induced one-way flows into treatment, to everyone. However, to preserve financial resources, in this scenario, the policymaker may wish to exclude individuals who are not likely to increase product usage, for example, because of low disease risks in their neighborhood.

Let $C:\mathcal{X}\times\mathcal{Z}\to\mathbb{R}_+$ be a user-chosen cost function that potentially depends on $\boldsymbol{\alpha}$. For example, $C(x,z)=|\boldsymbol{\alpha}(x,z)-z|$ is a direct measure of manipulation costs.\footnote{Depending on the context, additional costs can be embedded in the experimental design. For example, in the experiment documented in thornton2008demand, besides the monetary incentives for learning HIV results (the instrument), there were considerably high costs for testing, counseling/giving results, and selling condoms (see Table 12).} For encouragement rule $\boldsymbol{\alpha}$, I define its budget by aggregating the costs for individuals who actually take up the treatment: $B(\boldsymbol{\alpha})=E[C(X,Z)\cdot D(\boldsymbol{\alpha}(X,Z))]$. I consider settings in which the policymaker faces a harsh budget constraint such that the cost of implementing any encouragement rule cannot exceed $\kappa$.

remarkThe policymaker may only want to account for cost without imposing a fixed budget, which is the thought experiment considered in kitagawa2018should and chen2022personalized. In this case, one can redefine the social welfare criterion as $W(\boldsymbol{\alpha})-B(\boldsymbol{\alpha})$ to apply the analysis in Section (ref).

As in Section (ref), I specialize to binary encouragement rules when discussing the performance of statistical decision rules and denote the budget by $B(\pi)$ in place of $B(\boldsymbol{\alpha}^\pi)$. Given a class $\Pi$ of feasible encouragement rules,\footnote{I implicitly assume that there exists $\pi\in\Pi$ such that $B(\pi)\leq \kappa$.} the policymaker now solves a constrained optimization problem:

equation*[equation* omitted — 76 chars of source]

Let $\pi^*_{\mathrm{B}}$ denote the oracle solution. I follow sun2021empirical to introduce two desirable properties for statistical decision rules in the current setting: asymptotic optimality and asymptotic feasibility. Intuitively, with a large enough sample size, asymptotic optimality imposes that a statistical decision rule $\hat{\pi}$ is unlikely to achieve strictly lower welfare than $\pi^*_\mathrm{B}$, and asymptotic feasibility imposes that $\hat{\pi}$ is unlikely to strictly violate the budget constraint.

definitionA statistical decision rule $\hat{\pi}$ is asymptotically optimal if, for any $\epsilon>0$, \begin{equation*} \limsup_{n\to\infty}\Pr(W(\hat{\pi})-W(\pi^*_{\mathrm{B}})<-\epsilon)=0. \end{equation*} A statistical decision rule $\hat{\pi}$ is asymptotically feasible if, for any $\epsilon>0$, \begin{equation*} \limsup_{n\to\infty}\Pr(B(\hat{\pi})-\kappa>\epsilon)=0. \end{equation*}
remarkAsymptotic optimality and asymptotic feasibility are defined asymmetrically in sun2021empirical. On one hand, asymptotic optimality only requires the population welfare of a statistical decision rule to concentrate around the optimal value from below. On the other hand, asymptotic feasibility requires the statistical decision rule to satisfy the population budget constraint without any slackness and thus is extremely sensitive to sampling uncertainty. In consequence, sun2021empirical proves the negative result that no statistical decision rule can uniformly satisfy both properties over a sufficiently rich class of data generating processes. In contrast, after revising the definition of asymptotic feasibility to be symmetric with that of asymptotic optimality, I show that it is possible to construct a statistical decision rule that simultaneously achieves both properties.

Note that by ((ref)),

equation*[equation* omitted — 116 chars of source]

Define the budget-constrained EWM encouragement rule defined as a solution to the sample version of the population constrained optimization problem:

equation[equation omitted — 165 chars of source]

where

equation*[equation* omitted — 138 chars of source]

I set $\hat{\pi}_{\mathrm{BEWM}}=\emptyset$ if no $\pi\in\Pi$ satisfies $\hat{B}_n(\pi)\leq \kappa$. Theorem (ref) asserts that $\hat{\pi}_{\mathrm{BEWM}}$ satisfies both properties in Definition (ref). A proof is provided in (ref).

theoremSuppose that Assumptions (ref)--(ref) hold, and that $C(x,z)$ is uniformly bounded in $x$ and $z$. Then, $\hat{\pi}_{\mathrm{BEWM}}$ is asymptotically optimal and asymptotically feasible.

Empirical Application

In this section, I apply the feasible EWM encouragement rule and the budget-constrained EWM encouragement rule to provide guidance on how to encourage upper secondary schooling, using data from the third wave of the Indonesian Family Life Survey (IFLS) fielded from June through November 2000. carneiro2017average used this dataset to study the returns to upper secondary schooling in Indonesia. I follow carneiro2017average in restricting my sample to males aged 25--60 who are employed and who have non-missing reported wage and schooling information. This subsample consists of 2,104 individuals.\footnote{The subsample used in carneiro2017average does not contain the tuition fee variable, which plays a central role in my framework as the manipulatable instrument. Hence, I followed their descriptions to construct my subsample from raw data downloaded from the RAND Corporation website.}

I specify the relevant variables in my framework as follows. The outcome $Y$ is the log of hourly wages (in rupiah) constructed from self-reported monthly wages and hours worked per week. The treatment $D$ is an indicator of attendance of upper secondary school or higher, corresponding to 10 or more years of completed education. The first instrument $Z_1$ is the lowest fee per continuing student, in thousands of rupiah, among secondary schools in the community of current residence.\footnote{The term “community” refers to the lowest-level administrative division in Indonesia. A community can either be a desa (village) or a kelurahan (urban community).} The second instrument $Z_2$ is the distance, in kilometers, from the office of the community head of current residence to the nearest secondary school, which I define as the secondary school closest to the office of the community head.\footnote{The validity of an instrument constructed in this way can be controversial. Each individual's tuition and distance to school are based on their current residence rather than their residence at the time of the secondary schooling decision. Educated individuals may move to more urban areas with more schools and higher tuition fees. Nonetheless, I note that the instrumental variable independence assumption for unrestricted instruments has testable implications, which are the generalized instrumental inequalities proposed by kedagni2020generalized. Using their tests, I do not find evidence against the independence assumption between potential earnings and tuition fees, or between potential earnings and distance to school. } I treat $Z_1$ as manipulatable and $Z_2$ as not manipulatable. Collect $Z=(Z_1,Z_2)^\top$. The covariates $X$ include age, age squared, an indicator of rural residence, distance from the office of the community head of residence to the nearest health post, and indicators for religion, parental education, and the province of residence. Table (ref) in (ref) presents sample averages of these variables.

I focus on binary encouragement rules that manipulate $Z_1$ according to $\boldsymbol{\alpha}^\pi(x,z)=\pi(x,z)\cdot \alpha_1(x,z)+(1-\pi(x,z))\cdot \alpha_0(x,z)$. For the binary decision $\pi$, I consider the class of linear rules based on $(z_1,z_2)$:

equation*[equation* omitted — 170 chars of source]

I specify the manipulation function as $\alpha_1(x,z)=(z_1-a)\cdot 1\{z_1\geq a\}$ and $\alpha_0(x,z)=z_1$. Here, $\alpha_1$ describes a tuition subsidy of up to $a$ and $\alpha_0$ describes the status quo. The policymaker a priori chooses from $a\in\{2.5,22.25\}$, which correspond to the sample median and maximum of $Z_1$, respectively. I specify the cost function as $C(x,z)=|\boldsymbol{\alpha}^\pi(x,z)-z_1|$ and the budget constraint as $\kappa=0.28$, which is about one-tenth of the average hourly wage.

The fact that $Z_1$ has discrete support (with 56 distinct values) violates the support condition for nonparametric or semiparametric identification of the propensity score.\footnote{When $a=2.5$, only 20 out of the 35 support points of $\alpha_1(X,Z)$ lie in the support of $Z_1$. When $a=22.25$, $\alpha_1(X,Z)$ is identically equal to 0, which falls outside the support of $Z_1$.} Therefore, I estimate the propensity score from a logit regression of $D$ on $X$, $Z$, $Z_1\cdot Z_2$, and interactions between $Z$ and $X$.\footnote{This specification of propensity score is an adaptation of that considered by carneiro2017average and sasaki2021estimation, who use a single instrument $Z_2$.} Although all elements of $X$ are discrete, they together provide sufficient variation in the propensity score for the semiparametric estimation of the MTE.\footnote{The estimated propensity score takes 1,782 distinct values that almost cover the full unit interval.} I specify the conditional mean of $Y$ given $X$, $Z_2$, and $p(X,Z)$ as

equation*[equation* omitted — 98 chars of source]

where $G(\cdot)$ is an unknown function. By Corollary (ref), the social welfare criterion of encouragement rule $\pi$ is identified as $W(\pi)=E[Y]+E[\pi(X,Z)\cdot((p(X,\alpha_1(X,Z),Z_2)-p(X,Z))(X^\top,Z_2)(\beta_1-\beta_0)+G(p(X,\alpha_1(X,Z),Z_2))-G(p(X,Z)))]$. I use the double residual regression procedure of robinson1988root to estimate $(\beta_1,\beta_0)$. Given the estimators $\hat{p}(x,z)$ and $(\hat{\beta}_1,\hat{\beta}_0)$, I estimate $G(\cdot)$ using a nonparametric regression of the residual $Y-\hat{p}(X,Z)(X^\top,Z_2)\hat{\beta}_1-(1-\hat{p}(X,Z))(X^\top,Z_2)\hat{\beta}_0$ on $\hat{p}(X,Z)$. I use the locally linear regression throughout with a Gaussian kernel and a bandwidth of 0.06, which is determined by leave-one-out cross-validation.

I compute the feasible EWM encouragement rule $\hat{\pi}_{\mathrm{FEWM}}$ in ((ref)) and the budget-constrained EWM encouragement rule $\hat{\pi}_{\mathrm{BEWM}}$ in ((ref)) using the CPLEX mixed integer optimizer. Table (ref) presents point estimates of some key quantities of alternative encouragement rules. The first column reports the welfare gain, $W(\pi)-E[Y]$. The second column reports the share of eligible population for the tuition subsidy, $E[\pi(X,Z)]$. Based on the decomposition result in Corollary (ref), the third and fourth columns report the average change in treatment take-up, $E[p(X,\boldsymbol{\alpha}^\pi(X,Z),Z_2)]-E[p(X,Z)]$, and the PRTE, respectively. The former measures the proportion of individuals induced to enroll in or drop out of upper secondary school, and the latter measures the average change in the log of hourly wages among these individuals.

table[table omitted — 1,357 chars of source]

As can be seen from Table (ref), the seemingly favorable tuition subsidy has little effect on overall upper secondary school attendance when applied to everyone, resulting in a welfare gain of only a small magnitude. In contrast, the feasible EWM encouragement rule and the budget-constrained EWM encouragement rule achieve higher welfare gains by targeting a subpopulation with both a greater increase in treatment take-up and higher PRTE.

I plot the feasible EWM encouragement rule and the budget-constrained EWM encouragement rule in Panels A and B of Figure (ref), respectively. The shaded areas indicate the subpopulations to whom the tuition subsidy should be assigned. For both subsidy levels, the feasible EWM encouragement rule gives eligibility to individuals facing relatively high tuition fees and living relatively close to the nearest secondary school. The subpopulations targeted by the budget-constrained EWM encouragement rule shrink to the left. When the subsidy level $a$ is increased from 2.5 to 22.25, the budget-constrained EWM encouragement rule tends to prioritize individuals facing relatively low tuition fees.

figure[figure omitted — 172 chars of source]

Utilizing a decomposition of the social welfare criterion analogous to Corollary (ref), I offer a partial explanation of why the subpopulations indicated by the shaded areas in Figure (ref) are targeted. One can write

equation*[equation* omitted — 110 chars of source]

where

equation*[equation* omitted — 159 chars of source]

measures the average treatment effect among individuals with $(X,Z)=(x,z)$ who are induced to switch treatment status when going from the status quo to the tuition subsidy $\alpha_1$. Let $\operatorname{Med}(X)$ denote the sample median of $X$. I focus on the case of $a=22.25$, i.e., a full tuition waiver. Figure (ref) is based on point estimates of $p(x,z)$ and $\operatorname{PRTE}(x,z)$. Panel A displays the level sets of $(z_1,z_2)\mapsto p(\operatorname{Med}(X),\alpha_1(x,z),z_2)-p(\operatorname{Med}(X),z)$, namely the changes in treatment take-up for individuals with different values of $(Z_1,Z_2)$ and the median value of $X$. Panel B displays the level sets of $(z_1,z_2)\mapsto\operatorname{PRTE}(\operatorname{Med}(X),z)$.

figure[figure omitted — 599 chars of source]

From Panel A of Figure (ref), it can be seen that only individuals with low $Z_2$ are induced into treatment. From Panel B of Figure (ref), it can be seen that individuals induced into treatment have positive PRTE except for those with low values of $(Z_1,Z_2)$. Put together, absent budget constraints, individuals in the upper-left corner are prioritized for the full tuition waiver. Meanwhile, contour lines where $Z_1$ is high are steep in both panels, implying that individuals with higher $Z_1$ incur greater manipulation costs for the same amount of welfare gains. Consequently, the optimal policy under budget constraints, which trades off welfare gains against costs, gives up individuals with high $Z_1$.

Conclusion

In this paper, I propose a policy learning framework that allows for endogenous treatment selection by leveraging an instrumental variable. To deal with imperfect compliance when designing policies, I consider encouragement rules instead of treatment rules. To deal with failure of unconfoundedness when identifying the social welfare criterion, I incorporate the MTE function. Focusing on binary encouragement rules, I apply the representation of the social welfare criterion via the MTE to the EWM method and derive convergence rates of regret. I also consider extensions allowing for multiple instruments and budget constraints. I illustrate the EWM encouragement rule using data from the Indonesian Family Life Survey.

To be clear, the analysis in this paper critically relies on the point-identification of the social welfare criterion. The necessary support condition or parametric assumptions could be restrictive. An interesting avenue for future research is to incorporate approaches to policy learning under partial identification of policy parameters (e.g., russell2020policy, dadamo2022orthogonal, christensen2023optimal, yata2023optimal).