EconBase
← Back to paper

Regret Analysis in Threshold Policy Design

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,220 characters · 13 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Regret Analysis in Threshold Policy Design

titlepage\begin{abstract} Threshold policies are decision rules that assign treatments based on whether an observable characteristic exceeds a certain threshold. They are widespread across multiple domains, including welfare programs, taxation, and clinical medicine. This paper examines the problem of designing threshold policies using experimental data, when the goal is to maximize the population welfare. First, I characterize the regret -- a measure of policy optimality -- of the Empirical Welfare Maximizer (EWM) policy, popular in the literature. Next, I introduce the Smoothed Welfare Maximizer (SWM) policy, which improves the EWM's regret convergence rate under an additional smoothness condition. The two policies are compared by studying how differently their regrets depend on the population distribution, and investigating their finite sample performances through Monte Carlo simulations. In many contexts, the SWM policy guarantees larger welfare than the EWM. An empirical illustration demonstrates how the treatment recommendations of the two policies may differ in practice. \\ \\ \noindentKeywords: Threshold policies, heterogeneous treatment effects, statistical decision theory, randomized experiments. \\ \noindentJEL classification codes: C14, C44 \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

\doublespacing

Introduction

Treatments are rarely universally assigned. When their effects are heterogeneous across individuals, policymakers aim to target those who would benefit the most from specific interventions. Scholarships, for example, are awarded to students with high academic performance or financial need; tax credits are provided to companies engaged in research and development activities; medical treatments are prescribed to sick patients. Despite the potential complexity and multidimensionality of heterogeneous treatment effects, treatment eligibility criteria are often kept quite simple. This paper studies one of the most common of these simple assignment mechanisms: threshold policies, where the decision to assign the treatment is based on whether a scalar observable characteristic — referred to as the index — exceeds a specified threshold.

Threshold policies are ubiquitous, ranging across multiple domains. In welfare policies, they regulate the qualification for public health insurance programs through age card2008impact,shigeoka2014effect and anti-poverty programs through income crost2014aid. In taxation, they determine marginal rates through income brackets taylor2003corporation. In clinical medicine, the referral for liver transplantation depends on whether a composite of laboratory values obtained from blood tests is beyond a certain threshold kamath2007model. Even criminal offenses are defined through threshold policies: sanctions for Driving Under the Influence are based on whether the Blood Alcohol Content exceeds specific values.

Economists frequently study the outcomes of threshold policies, with the regression discontinuity design (RDD) being a widely used tool for causal inference. The RDD focuses on an ex-post evaluation of treatment effects at discontinuity points. In this paper, my perspective is different: I consider the ex-ante problem faced by a policymaker seeking to implement a threshold policy and interested in maximizing the average social welfare, targeting individuals who would benefit from the treatment. Experimental data are available: how should they be used to design the threshold policy for the population?

Answering this question requires defining a criterion by which policies are evaluated. Since the performance of a policy depends on the unknown data distribution, the policymaker searches for a policy that behaves uniformly well across a specified family of data distributions (the state space). The regret of a policy is the (possibly) random difference between the maximum achievable welfare and the welfare it generates in the population. Policies can be evaluated considering their maximum expected regret manski2004statistical, hirano2009asymptotics, kitagawa2018should, athey2021policy, mbakop2021model, manski2023probabilistic, or other worst-case statistics of the regret distribution manski2023statistical, kitagawa2022treatment. Once the criterion has been established, optimal policy learning aims to pinpoint the policy that minimizes it. Rather than directly tackling the functional minimization problem, following the literature, I consider candidate threshold policy functions and characterize some properties of their regret.

The first contribution of this paper is to show how to derive the asymptotic distribution of the regret for a given threshold policy. The underlying intuition is simple: threshold policies use sample data to choose the threshold, which is hence a random variable with a certain asymptotic behavior. A Taylor expansion establishes a map between the regret of a policy and its threshold, allowing one to characterize the asymptotic distribution of the regret through the asymptotic behavior of the threshold. This shifts the problem to characterizing the asymptotic distribution of the threshold estimator, simplifying the analysis as threshold estimators can be studied with common econometric tools.

I start considering the Empirical Welfare Maximizer (EWM) policy studied by kitagawa2018should. They derive uniform bounds for the expected regret of the policy for various policy classes, where the policy class impacts the findings only in terms of its VC dimensionality. My approach is more specific, considering only threshold policies, but also more informative: leveraging the knowledge of the policy class, I characterize the asymptotic distribution of the regret. As mentioned above, this requires the derivation of the asymptotic distribution of the threshold for the EWM policy, which is non-standard: it exhibits the “cube root asymptotics” behavior studied in kim1990cube. The convergence rate is $n^{\frac{1}{3}}$, and the asymptotic distribution is of chernoff1964estimation type. The non-standard behavior and the unusual convergence rate are due to the discontinuity in the objective function and are reflected in the asymptotic distribution of the regret, and in its $n^{\frac{2}{3}}$ pointwise convergence rate.

My second contribution is hence the introduction of a novel threshold policy, the Smoothed Welfare Maximizer (SWM) policy. This approach modifies the Empirical Welfare Maximizer (EWM) policy by replacing the indicator function in its objective function with a smooth kernel. Under certain smoothness regularity assumptions, the threshold estimator for the SWM policy is asymptotically normal, and its regret achieves a pointwise convergence rate of $n^{\frac{4}{5}}$. While this paper does not establish the optimality of this rate -- meaning it does not determine whether the SWM policy attains the fastest possible convergence rate under the given regularity conditions -- my results imply that, under the additional assumption, the SWM policy achieves a faster pointwise convergence rate than the EWM policy.

Building on these asymptotic results, I extend the comparison of the regrets with the EWM and the SWM policies beyond their convergence rates. My findings allow to compare the asymptotic distributions and investigate how differently they depend on the data distribution; theoretical results are helpful to inform and guide the Monte Carlo simulations, which confirm that the asymptotic results approximate the actual finite sample behaviors. Notably, the simulations confirm that the SWM policy may guarantee lower expected regret in finite samples.

To demonstrate the practical differences between the two policies, I present an empirical illustration considering a job-training treatment. In that context, the SWM threshold policy would recommend treating 66.2% of unemployed workers, as opposed to 63.6% with the EWM policy. This difference of almost 3 percentage points is economically non-negligible.

Related Literature

This paper relates to the statistical decision theory literature studying the problem of policy assignment with covariates manski2004statistical,stoye2012minimax,kitagawa2018should,athey2021policy,mbakop2021model,sun2021treatment,sun2021empirical,viviano2023fair. My setting is mainly related to kitagawa2018should and athey2021policy, with some notable differences. kitagawa2018should study the EWM policy for policy classes with finite VC dimension. They derive finite sample bounds for the expected regret without relying on smoothness assumptions. athey2021policy consider a double robust version of the EWM and allow for observational data. Under smoothness assumptions analogous to mine, they derive asymptotic bounds for the expected regret for policy classes with finite VC dimensions. Conversely, results in this paper apply exclusively to threshold policies, relying on a combination of the assumptions in kitagawa2018should and athey2021policy. The narrower focus allows for more comprehensive results: I derive the asymptotic distribution of the regret, rather than providing some bounds for the expected regret. A critical distinction lies in the different nature of the convergence rates. My results are valid pointwise, derived by leveraging additional assumptions on the data distribution. As a result, the rates I obtain for the EWM and SWM policies are faster than the $\sqrt{n}$ rate reported as optimal by kitagawa2018should and athey2021policy. Their $\sqrt{n}$ rate is, in fact, uniformly valid for a broader family of data distributions, including extreme cases (e.g., where conditional ATE is flat at the threshold) that are excluded from my analysis. Their uniform results may be viewed as a benchmark: when more structure is imposed on the problem and certain data distributions are excluded, the rates can be improved.

Optimal policy learning finds its empirical counterpart in the literature dealing with targeting, especially common in development economics. Recent studies rely on experimental evidence to decide who to treat in a data-driven way hussam2022targeting,aiken2022machine, even if haushofer2022targeting pointed out the need for a more formalized approach to the targeting decision problem. The availability for the policymaker of appropriate tools to use the data in the decision process is probably necessary to guarantee a broader adoption of data-driven targeting strategies. Focusing on threshold policies, this paper explicitly formulates the decision problem, introduces implementable policies (the EWM and the SWM policy), and compares their asymptotic properties.

Turning to the threshold estimators, I already mentioned that the EWM policy exhibits the cube root of $n$ asymptotics studied by kim1990cube, distinctive of several estimators in different contexts. Noteworthy examples are the maximum score estimator in choice models manski1975maximum, the split point estimator in decision trees banerjee2007confidence, and the risk minimizer in classification problems mohammadi2005asymptotics, among others. Specific to my analysis is the emergence of the cube root asymptotic within a causal inference problem relying on the potential outcomes model, which is then mirrored in the regret's asymptotic distribution.

Addressing the cube root problem by smoothing the indicator in the objective function aligns closely with the strategy proposed by horowitz1992smoothed for studying the asymptotic behavior of the maximum score estimator. Objective functions are nonetheless different, and in my context, I derive the asymptotic distribution for both the unsmoothed (EWM) and the smoothed (SWM) policies. This is convenient, as it allows me to compare not only the convergence rates but also the entire asymptotic distributions of the estimators and their regrets and study the asymptotic approximations in Monte Carlo simulations.

The rest of the paper is structured as follows. Section (ref) introduces the problem and outlines my analytical approach. Section (ref) derives formal results for the asymptotic distribution of the EWM and SWM policies and their regrets. In Section (ref), I investigate finite sample performance of the EWM and SWM policies through Monte Carlo simulations, while in Section (ref) I consider the analysis of experimental data from the National Job Training Partnership Act (JTPA) Study to compare the practical implications of the policies. Section (ref) concludes.

Overview of the Problem

I consider the problem of a policymaker who wants to implement a binary treatment in a population of interest. Individuals are characterized by a vector of observable characteristics $\textbf{X} \in \mathbb{R}^d$, on which the policymaker bases the treatment assignment choice. A policy is hence a map $\pi(\textbf{x}): \mathbb{R}^d \rightarrow \{0,1\}$, from observable characteristics to the binary treatment status. The policymaker is utilitarian: its goal is to maximize the average welfare of the population. Indicating by $Y_1$ and $Y_0$ the potential outcomes with and without the treatment, population average welfare generated by a policy $\pi$ can be written as

gather[gather omitted — 89 chars of source]

When treatment effects are heterogeneous, the same treatment can have opposite average effects across individuals with different $\textbf{X}$'s. For this reason, the policy assignment may vary with $\textbf{X}$: the policymaker wants to target only those who benefit from being treated, to maximize the average welfare.

The policy learning literature has considered several classes $\Pi$ of policy functions, such as linear eligibility indexes, decision trees, or monotone rules, discussed in kitagawa2018should,athey2021policy, mbakop2021model. This paper focuses on threshold policies, a specific class of policy functions that can be represented as

gather[gather omitted — 69 chars of source]

The threshold policy assigns the treatment whenever the scalar index $X \in \mathbb{R}$, one of the observable characteristics, exceeds a threshold $t$, the parameter to be chosen.

Threshold policies are widespread: they regulate, beyond others, organ transplants kamath2007model, taxation taylor2003corporation, and access to social welfare programs card2008impact,crost2014aid. Their key advantage seems to be simplicity: threshold policies are easy for eligible individuals to understand, simple for policymakers to implement and monitor, and transparent, with clearly defined eligibility criteria -- unlike more opaque black-box algorithms. These factors often justify the use of threshold policies even when a more structured alternative policy class may theoretically deliver higher welfare. In practice, these alternatives require additional resources for implementation, adoption, and monitoring -- potentially offsetting the welfare gains. Modeling this trade-off goes beyond the scope of this paper, where the restriction to the threshold policy class is taken as given and should not be interpreted as an endorsement of threshold policies. Nonetheless, it is worth noting that if the conditional average treatment effect (CATE) is monotone in $X$ and exhibits sign heterogeneity, as is often the case in applications, then the threshold policy is optimal among all policies that use only $X$ to assign the treatment.

I will focus on the case when the index $X$ is chosen before the experiment. Population welfare depends only on threshold $t$, and can be written as

gather[gather omitted — 97 chars of source]

Choosing the policy is equivalent to choosing the threshold. If the joint distribution of $Y_1$, $Y_0$, and $X$ were known, the policymaker would implement the policy with threshold $t^*$ defined as:

gather[gather omitted — 139 chars of source]

which would guarantee the maximum achievable welfare $W(t^*)$.

The problem described in equation (ref) is infeasible since the joint distribution of $Y_1$, $Y_0$, and $X$ is unknown. The policymaker observes an experimental sample $Z=\{Z_i\}_{i=1}^n=\{Y_i,D_i,X_i\}$, where $Y$ is the outcome of interest, $D$ the randomly assigned treatment status, and $X$ the policy index. Experimental data, which allows to identify the conditional average treatment effect, are used to learn the threshold policy $\hat{t}_n = \hat{t}_n(Z)$, function of the observed sample.

Statistical decision theory deals with the problem of choosing the map $\hat{t}_n$. First, it is necessary to specify the decision problem the policymaker faces. For any threshold policy $\hat{t}_n$, define the regret $\mathcal{R}(\hat{t}_n)$:

gather[gather omitted — 65 chars of source]

a measure of welfare loss indicating the suboptimality of policy $\hat{t}_n$. The regret depends on the unknown data distribution: the policymaker specifies a state space, and searches for a policy that does well uniformly for all the data distributions in the state space. Following manski2004statistical, statistical decision theory has mainly focused on the problem of minimizing the maximum expected regret, looking for a policy $\hat{t}_n$ that does uniformly well on average across repeated samples.

Directly solving the constrained minimization problem of the functional $\sup \mathbb{E}[\mathcal{R}(\hat{t}_n)]$ is impractical: the literature instead focuses on considering a specific policy map and studying its properties, for example showing its rate optimality, through finite sample valid kitagawa2018should or asymptotic athey2021policy arguments. Following this approach, I characterize and compare some properties for the regret of two different threshold policies, the Empirical Welfare Maximizer (EWM) policy, commonly studied in the literature, and the novel Smoothed Welfare Maximizer (SWM) policy.

kitagawa2018should derive finite sample bounds for the expected regret of the EWM policy for a wide range of policy function classes. In their results, the policy class $\Pi$ affects the bounds only through its VC dimension, and the knowledge of $\Pi$ is not further exploited. Conversely, I leverage the additional structure from the knowledge of the policy class and characterize the asymptotic distribution of the regret for the EWM and the SWM threshold policies, comparing how their regrets depend on the data distribution. My results could hence be of interest also when decision problems not involving the expected regret are considered, as in manski2023statistical and kitagawa2022treatment: I characterize the asymptotic behavior of regret quantiles, and the asymptotic distributions can be used to simulate expectations of their non-linear functions.

To derive my results, I take advantage of the link between a threshold policy function $\hat{t}_n$ and its regret $\mathcal{R}(\hat{t}_n)$. Let $\{r_n\}$ be a sequence such that $r_n \rightarrow \infty$ for $n \rightarrow \infty$, and suppose that $r_n (\hat{t}_n - t^*)$ converges to a non degenerate limiting distribution, i.e $(\hat{t}_n - t^*) = O_p(r_n^{-1})$.

Assume function $W(t)$ to be twice continuously differentiable, and consider its second-order Taylor expansion around $t^*$:

gather*[gather* omitted — 153 chars of source]

where $|\tilde{t}-t^*| \leq |\hat{t}_n - t^*|$, and $W'(t^*)=0$ by optimality of $t^*$. The previous equation can be written as

gather[gather omitted — 148 chars of source]

establishing a relationship between the convergence rates of $\hat{t}_n$ and $\mathcal{R}(\hat{t}_n)$, and between their asymptotic distributions. Equation (ref) therefore shows how the rate of convergence and the asymptotic distribution of regret $\mathcal{R}(\hat{t}_n)$ can be studied through the rate of convergence and the asymptotic distribution of policy $\hat{t}_n$. In the next section, I consider the EWM policy $\hat{t}^e_n$ and the SWM policy $\hat{t}^s_n$: through their asymptotic behaviors, I characterize the asymptotic distributions of their regrets $\mathcal{R}(\hat{t}^e_n)$ and $\mathcal{R}(\hat{t}^s_n)$.

rem{\normalfont (Ceteris Paribus Optimality)} Following the literature in statistical decision theory, I assume that the experiment is conducted in a population with the same distribution as the one where the policy will be implemented. This implicitly assumes that individuals in the target population do not change their behavior in response to the policy -- for instance, by altering their covariates to gain access to the treatment. In some contexts, this assumption may be unrealistic. Existing empirical studies on threshold policies highlight this issue: in regression discontinuity design, manipulation tests are specifically aimed at detecting such reactions to the policy. If manipulation occurs, it invalidates the optimality of the policy estimated in the experiment.

Formal Results

Let $Y_0$ and $Y_1$ be scalar potential outcomes, $D$ the binary treatment assignment in the experiment, and $X$ the observable index. $\{Y_0,Y_1, D, X \}$ are random variables distributed according to the distribution $P$. They satisfy the following assumptions, which guarantee the identification of the optimal threshold:

ass{\normalfont (Identification)} \begin{enumerate}[label=1.\arabic*] • {\normalfont (No interference)} Observed outcome $Y$ is related with potential outcomes by the expression $Y= DY_1 + (1-D)Y_0$. • {\normalfont (Unconfoundedness)} Distribution $P$ satisfies $D \perp\!\!\!\perp (Y_0,Y_1) |X$. • {\normalfont (Overlap)} Propensity score $p(x)=\mathbb{E}[D|X=x]$ is assumed to be known and such that $p(x) \in (\eta,1-\eta)$, for some $\eta \in (0,0.5)$. • {\normalfont (Joint distribution)} Potential outcomes $(Y_0, Y_1)$ and index $X$ are continuous random variables with joint probability density function $\varphi (y_0,y_1,x)$, and marginal densities $\varphi_0$, $\varphi_1$, and $f_x$ respectively. Expectations $\mathbb{E}[Y_0|x]$ and $\mathbb{E}[Y_1|x]$, for $x$ in the support of $X$, exist. \end{enumerate}

Assumptions (ref), (ref), and (ref) are standard assumptions in many causal models. Assumption (ref) requires the outcome of each unit to depend only on their treatment status, excluding spillover effects. Assumption (ref) requires random assignment of the treatment, conditionally on $X$. Assumption (ref) requires that, for any value of $X$, there is a positive probability of observing both treated and untreated units. Probabilities of being assigned to the treatment may vary with $X$, allowing for stratified experiments.

Assumption (ref) specifies the focus on continuous outcome and index. While it would be possible to accommodate discrete $Y_0$ and $Y_1$, maintaining the continuity of $X$ remains essential. The arguments developed in this paper, in fact, are not valid for a discrete index: my focus is on studying optimal threshold policies in contexts where the probability of observing any value on the support of the index $X$ is zero, and the threshold must be chosen from a continuum of possibilities.

Under Assumption (ref), optimal policy $t^*$ defined in (ref) can be written as

align[align omitted — 268 chars of source]

and is hence identified. This standard result specifies under which conditions an experiment allows to identify $t^*$.

Empirical Welfare Maximizer Policy

Policymaker observes an i.i.d. random sample $Z=\{Y_i,D_i,X_i\}$ of size $n$ from $P$, and considers the Empirical Welfare Maximizer policy $\hat{t}^e_n$, the sample analog of $t^*$ in equation (ref)\footnote{The objective function is piecewise constant, so solving the optimization problem requires evaluating the function $n+1$ times. The solution is the convex set of points in $\mathbb{R}$ that achieve the maximum of these values.}:

gather[gather omitted — 219 chars of source]

Policy $\hat{t}^e_n$ can be seen as an extremum estimator, maximizer of a function not continuous in $t$.

Consistency of $\hat{t}^e_n$

First, I will prove that $\hat{t}^e_n$ consistently estimates the optimal threshold $t^*$, implying that $\mathcal{R}(\hat{t}^e_n) \rightarrow ^p 0$. To prove this result, I need the following assumptions on the data distribution.

ass{\normalfont (Consistency)} \begin{enumerate}[label=2.\arabic*] • {\normalfont (Maximizer $t^*$)} Maximizer $t^* \in \mathcal{T} $ of $\mathbb{E}[(Y_1-Y_0) \mathbf{1}\{X> t\}]$ exists and is unique. It is an interior point of the compact parameter space $\mathcal{T} \subseteq \mathbb{R}$. • {\normalfont (Square integrability)} Conditional expectations $\mathbb{E}[Y_0^2|X]$ and $\mathbb{E}[Y_1^2|X]$ exist. • {\normalfont (Smoothness)} In a neighbourhood of $t^*$, density $f_x(x)$ is positive, and function $\mathbb{E}[(Y_1-Y_0) \mathbf{1}\{X> t\}]$ is at least $s$-times continuously differentiable in $t$. \end{enumerate}

By requiring the existence of the optimal threshold in the interior of the parameter space, Assumption (ref) is assuming heterogeneity in the sign of the conditional average treatment effect $\mathbb{E}[Y_1 - Y_0 | X]$. It is because of this heterogeneity that the policymaker implements the threshold policy, targeting groups that would benefit from being treated. The assumption neither excludes the multiplicity of local maxima, as long as the global one is unique, nor excludes unbounded support for $X$, but requires the parameter space to be compact. Uniqueness of the maximizer is not required in the standard analysis of the EWM policy, where partial identification of the optimal policy is allowed. A sufficient condition for Assumption (ref), easy to interpret and plausible in many applications, is that the conditional average treatment effect has negative and positive values, and crosses zero exactly once.

Assumption (ref) requires that the conditional potential outcomes have finite second moments and is satisfied when $Y$ is assumed to be bounded (as in kitagawa2018should).

Assumption (ref) will be used with increasing values of $s$ to prove different results. To prove consistency, it needs to hold for $s=0$, requiring the continuity of the objective function $W(t)$ in a neighborhood of $t^*$. The derivative of $\mathbb{E}[(Y_1-Y_0) \mathbf{1}\{X> t\}]$ with respect to $t$ is equal to $-f_x(t) \tau(t)$, where $\tau(x) = \mathbb{E}[Y_1-Y_0 | X=x]$ is the conditional average treatment effect. Assumption (ref) with $s \geq 1$ hence requires smoothness of $f_x(x)$ and $\tau(x)$, in a neighborhood of $t^*$.

The following theorem proves the consistency of $\hat{t}^e_n$ for $t^*$.

thmConsider the EWM policy $\hat{t}^e_n$ defined in equation (ref) and the optimal policy $t^*$ defined in equation (ref). Under Assumptions (ref) and (ref) (with $s=0$), \begin{gather*} \hat{t}^e_n \rightarrow^{a.s.} t^* \end{gather*} i.e. $\hat{t}^e_n$ is a strongly consistent estimator for $t^*$.

Asymptotic Distribution for $\hat{t}^e_n$

The fact that $\hat{t}^e_n$ is the maximizer of a function not continuous in $t$ directly affects the convergence rate and the asymptotic distribution. The EWM policy $\hat{t}^e_n$ exhibits the “cube root asymptotics” behavior studied in kim1990cube, the same as, beyond others, the maximum score estimator manski1975maximum, and the split point estimator in decision trees banerjee2007confidence.

The limiting distribution is not Gaussian, and its derivation requires two additional regularity conditions on $P$:

ass{\normalfont (Asymptotic Distribution)} \begin{enumerate}[label=3.\arabic*] • {\normalfont (Non-flat $\tau(X)$)} In a neighbourhood of $t^*$, the derivative of the conditional average treatment effect, $\frac{\partial \mathbb{E}\left[Y_1 - Y_0 | X \right]}{\partial X}$, is non-zero. • {\normalfont (Tail condition)} Let $\varphi_1$ and $\varphi_0$ be the probability density functions of $Y_1$ and $Y_0$. Assume that, as $|y| \rightarrow \infty$, $\varphi_1(y) = o(|y|^{-(4+\delta)})$ and $\varphi_0(y) = o(|y|^{-(4+\delta)})$, for $\delta > 0$. \end{enumerate}

Assumption (ref) requires that, close to the maximizer $t^*$, the conditional average treatment effect function $\tau(X)$ is not flat. If $\tau(X)$ were flat in the neighborhood of $t^*$, it would be harder for the estimator to find the exact maximizer, leading to a slower rate of convergence. The excluded flat $\tau(X)$ corresponds to a situation where estimating the threshold precisely is less critical, as the CATE remains zero even in a neighborhood of the optimal threshold.

Assumption (ref) requires that the tails of the distributions of the potential outcomes are not too fat. It is generally satisfied by any bounded distribution, and by distributions in the exponential family, while is violated, for example, by the Student's t-distribution with fewer than four degrees of freedom.

The following theorem gives the asymptotic distribution of $\hat{t}^e_n$.

thmConsider the EWM policy $\hat{t}^e_n$ defined in equation (ref) and the optimal policy $t^*$ defined in equation (ref). Under Assumptions (ref), (ref) (with $s=2$), and (ref), as $n\rightarrow \infty$, \begin{gather} n^{1 / 3}\left(\hat{t}^e_n-t^*\right) \rightarrow^d (2\sqrt{K}/H)^{\frac{2}{3}}\mathop{\rm arg max}\limits_r \left(B(r) - r^2 \right) \end{gather} where $B(r)$ is the two-sided standard Brownian motion process, and $K$ and $H$ are \begin{align*} K =& f_x(t^*) \left(\frac{1}{p(t^*)} \mathbb{E}[Y_1^2|X = t^*] + \frac{1}{1-p(t^*)} \mathbb{E}[Y_0^2|X = t^*] \right) \\ H =& f_x(t^*) \left(\frac{\partial \mathbb{E}\left[Y_1 - Y_0 | X = t^* \right]}{\partial X} \right). \end{align*}

The limiting distribution of $n^{1 / 3}\left(\hat{t}^e_n-t^*\right)$ is of Chernoff type chernoff1964estimation. The Chernoff's distribution is the probability distribution of the random variable $\mathop{\rm arg~max}\limits_r B(r) - r^2$, where $B(r)$ is the two-sided standard Brownian motion process. The process $B(r) - r^2$ can be simulated, and the distribution of $\mathop{\rm arg~max}\limits_r B(r) - r^2$ numerically studied. groeneboom2001computing report values for selected quantiles.

It's worth noticing how the variance of $\hat{t}^e_n$ depends on the data distribution. $K$ and $H$ are functions of the density of $X$, the variance of the potential outcomes, and the derivative of the CATE at $t^*$. The optimal threshold is estimated with more precision when more data around the optimal threshold are available (larger density), when the treatment effect changes more rapidly (larger derivative of CATE), and when the outcomes have less variability.

Results in Theorem (ref) can be used to derive asymptotic valid confidence intervals for $\hat{t}^e_n$, as discussed in Appendix (ref). More interestingly, they can be combined with Equation (ref) to characterize the asymptotic distribution of the regret $\mathcal{R}(\hat{t}^e_n)$, as derived in the following corollary.

corollaryThe asymptotic distribution of regret $\mathcal{R}(\hat{t}^e_n)$ is: \begin{gather*} n^{\frac{2}{3}} \mathcal{R}(\hat{t}^e_n) \rightarrow^d \left( \frac{2K^{2}}{H} \right)^\frac{1}{3} \left( \mathop{\rm arg max}\limits_r B(r) - r^2 \right)^2. \end{gather*} The expected value of the asymptotic distribution is $K^\frac{2}{3} H^{-\frac{1}{3}} C^e$, where $$C^e= \sqrt[3]{2} \mathbb{E}\left[\left( \mathop{\rm arg~max}\limits_r B(r) - r^2 \right)^2\right]$$ is a constant not dependent on $P$.

For the regret of the EWM policy, Corollary (ref) establishes a $n^{\frac{2}{3}}$ rate, faster than the $\sqrt{n}$ rate found to be the optimal for the EWM expected regret kitagawa2018should. It is essential to highlight the differences between the two results: Corollary (ref) is about pointwise convergence in distribution, while the main results by kitagawa2018should establish a uniform rate for the expected regret. The family of distributions they consider may violate the assumptions in Theorem (ref), for example including cases where the CATE is flat at the optimal threshold. In contrast, Corollary (ref) is derived for distributions that satisfy Assumptions (ref), (ref) and (ref), which imply $H \neq 0$.

kitagawa2018should also discuss how additional assumptions on the distribution $P$ can lead to a faster convergence rate for the regret. They show that if a certain margin condition on the data distribution holds, the convergence rate can improve. When the CATE function $\tau(X)$ is non-flat, the regret achieves a uniform convergence rate of $n^{\frac{2}{3}}$. Although the two rates are the same and rely on a similar condition on the CATE, they are derived from different, non-nested sets of assumptions. For instance, kitagawa2018should assume that the policy class is correctly specified and $Y$ is bounded, but don't require the CATE to be smooth. The fact that my pointwise convergence rate coincides with their uniform rate when all assumptions are met suggests that my rate for the EWM threshold policy is also uniformly optimal.

The result in Corollary (ref) does not imply convergence in the mean, and the expected value of the asymptotic distribution is presented as a summary statistic -- a measure of the location of the asymptotic distribution.

Smoothed Welfare Maximizer Policy

Corollary (ref) shows how the cube root of $n$ convergence rate of the EWM policy directly impacts the convergence rate of its regret. In this section, I propose an alternative threshold policy, the Smoothed Welfare Maximizer policy, that achieves a faster rate of convergence and hence guarantees a faster rate of convergence for its regret. My approach exploits some additional smoothness assumptions on the distribution $P$: Corollary (ref) holds when $f_x(x)$ and $\tau(x)$ are assumed to be at least once differentiable; if they are at least twice differentiable, the SWM policy guarantees a $n^\frac{4}{5}$ convergence rate for the regret. Note that asking the density of the index and the conditional average treatment effect to be twice differentiable seems plausible for many applications. In the context of policy learning, it is, for example, assumed by athey2021policy to derive their results\footnote{This assumption is implied by the high-level assumptions made in the paper, as the authors discuss in footnote 15.}.

My approach involves smoothing the objective function in (ref), in the same spirit as the smoothed maximum score estimator proposed by horowitz1992smoothed to deal with inference for the maximum score estimator manski1975maximum. The Smoothed Welfare Maximizer (SWM) policy $\hat{t}^s_n$ is defined as\footnote{The objective function is continuous and smooth in $t$, which allows for the use of gradient descent algorithms. However, the function is not concave and may have many local maxima. Annealing algorithms can be employed to find the global maximum.}:

gather[gather omitted — 268 chars of source]

where $\sigma_n$ is a sequence of positive real numbers such that $\lim_{n \rightarrow \infty} \sigma_n = 0$, and the function $k(\cdot)$ satisfies:

ass{\normalfont (Kernel function)} Kernel function $k(\cdot): \mathbb{R} \rightarrow \mathbb{R}$ is continuous, bounded, and with limits $\lim_{x \rightarrow - \infty} k(x) = 0$ and $\lim_{x \rightarrow \infty} k(x) = 1$.

In practice, the indicator function found in $\hat{t}^e_n$ is here substituted by a smooth function $k(\cdot)$ with the same limiting behavior, which guarantees the differentiability of the objective function. The bandwidth $\sigma_n$, decreasing with the sample size, ensures that when $n \to \infty$ the policy converges to the optimal one, as proved in the next section.

Consistency of $\hat{t}^s_n$

I start showing consistency of $\hat{t}^s_n$ for $t^*$, which implies $\mathcal{R}(\hat{t}^s_n) \rightarrow^p 0$.

thmConsider the SWM policy $\hat{t}^s_n$ defined in equation (ref) and the optimal policy $t^*$ defined in equation (ref). Under Assumptions (ref), (ref) (with $s=0$), and (ref), as $n \rightarrow \infty$, \begin{gather*} \hat{t}^s_n \rightarrow^{a.s.} t^* \end{gather*} i.e. $\hat{t}^s_n$ is a strongly consistent estimator for $t^*$.

Theorems (ref) and (ref) are analogous: they rely on the same assumptions on the data (Assumptions (ref) and (ref)) to prove the consistency of $\hat{t}^e_n$ and $\hat{t}^s_n$. Where the two policies differ is in the asymptotic distributions: smoothness in the objective function for $\hat{t}^s_n$ guarantees asymptotic normality, but also introduces a bias, since the bandwidth $\sigma_n$ equals zero only in the limit, which emerges in the limiting distribution.

Asymptotic Distribution for $\hat{t}^s_n$

Deriving this asymptotic behavior of $\hat{t}^s_n$ requires an additional assumption on the rate of bandwidth $\sigma_n$ and the kernel function $k$. Since both are chosen by the policymaker, the assumption is not a restriction on the data but a condition on properly picking $\sigma_n$ and $k$.

ass{\normalfont (Bandwidth and kernel)} \begin{enumerate}[label=5.\arabic*] • {\normalfont (Rate of $\sigma_n$)} $\frac{\log n}{n \sigma_n^4} \rightarrow 0$ as $n\rightarrow \infty$. • {\normalfont (Kernel function)} Kernel function $k(\cdot): \mathbb{R} \rightarrow \mathbb{R}$ satisfies Assumption (ref) and the following: \begin{itemize} • $k(\cdot)$ is twice differentiable, with uniformly bounded derivatives $k'$ and $k''$. • $\int k'(x)^4 dx$, $\int k''(x)^2 dx$, and $\int |x^2 k''(x)| dx$ are finite. • For some integer $h \geq 2$ and each integer $i\in[1,h]$, $\int |x^i k'(x)| dx =0$ for \\ $i<h$ and $\int |x^h k'(x)| dx =d \neq 0$, with $d$ finite. • For any integer $i \in [0,h]$, any $\eta >0$, and any sequence $\sigma_n \rightarrow 0$, \\ $\lim_{n\rightarrow \infty} \sigma_n^{i-h} \int_{|\sigma_n x| > \eta} | x^i k'(x)| dx = 0$, and $\lim_{n\rightarrow \infty} \sigma_n^{-1} \int_{|\sigma_n x| > \eta} | k''(x)| dx = 0$. • $\int x k''(x) dx = 1$, $\lim_{n\rightarrow \infty} \int_{|\sigma_n x| > \eta} | x k''(x)| dx = 0$. \end{itemize} \end{enumerate}

An example of a function $k$ satisfying Assumption (ref) with $h=2$ is the cumulative distribution function of the standard normal distribution.

I can now derive the asymptotic distribution of $\hat{t}^s_n$.

thmConsider the SWM policy $\hat{t}^s_n$ defined in equation (ref) and the optimal policy $t^*$ defined in equation (ref). Under Assumptions (ref), (ref) (with $s=h + 1$ for some $h\geq 2$), (ref), and (ref), as $n \rightarrow \infty$: \begin{enumerate} • if $n \sigma_n^{2h + 1} \rightarrow \infty$, $$\sigma_n^{-h}(\hat{t}^s_n - t^*) \rightarrow^p H^{-1}A;$$ • if $n \sigma_n^{2h + 1} \rightarrow \lambda < \infty$, $$(n\sigma_n)^{\frac{1}{2}}(\hat{t}^s_n - t^*) \rightarrow^d \mathcal{N}(\lambda^{\frac{1}{2}}H^{-1}A, H^{-2}\alpha_2 K);$$ \end{enumerate} where $A$, $\alpha_1$, and $\alpha_2$ are: \begin{align} A =& -\frac{1}{h!} \alpha_1 \int_y \left(Y_1 - Y_0 \right) \varphi^{h}_x(y,t^*) d y \\ \alpha_1 =& \int_\zeta \zeta^h k'\left(\zeta \right) d \zeta \\ \alpha_2 =& \int_\zeta k'\left(\zeta \right)^2 d \zeta. \end{align}

The asymptotic distribution of $(n\sigma_n)^{\frac{1}{2}}(\hat{t}^s_n - t^*)$ is normal, centered at the asymptotic bias $\lambda^{\frac{1}{2}}H^{-1}A$ introduced by the smoothing function, which exploits local information giving non-zero weights to treated units in the untreated region and vice versa. The bias and the variance of the distribution depend on the population distribution through $K$ and $H$, as for the EWM policy, and also through $A$, a new term that determines the bias. In the definition of $A$, $\varphi^{h}_x$ is the $h$ derivative with respect to $x$ of $\varphi(y_0,y_1,x)$, the joint density distribution of $Y_0$, $Y_1$, and $X$: the integral in the expression for $A$ is the $h$-derivative of $f_x(X) \tau(X)$ computed in $X=t^*$, whose existence is guaranteed by Assumption (ref) with $s=h+1$. $\alpha_1$ and $\alpha_2$ depends only on kernel function $k$, and are hence known.

Theorems (ref) and (ref) differ in the smoothness requirements imposed by Assumption (ref). Theorem (ref) requires the function $\mathbb{E}[(Y_1-Y_0) \mathbf{1}\{X> t\}]$ to be at least twice differentiable, while Theorem (ref) requires three derivatives. The additional smoothness condition is necessary because certain steps in the proof rely on a Taylor expansion of the joint distribution of $Y$ and $X$ which requires $s \geq 3$ to exist. When $s=2$, the asymptotic distribution derived in Theorem (ref) for $\frac{\partial \hat{S}_n(\hat{t}^s_n,\sigma_n)}{\partial^2 t} (n\sigma_n)^{\frac{1}{2}}(\hat{t}^s_n - t^*)$ no longer holds, as $\frac{\partial \hat{S}_n(\hat{t}^s_n,\sigma_n)}{\partial^2 t}$, the second derivative of the objective function in the SWM policy definition (Equation (ref)), may not have a bounded limiting distribution. However, if such a bounded limiting distribution, denoted by $\tilde{H}$, exists, then $(n\sigma_n)^{\frac{1}{2}}(\hat{t}^s_n - t^*)$ remains $O_p(1)$ even when $s=2$, with an asymptotic distribution that depends on the unknown term $\tilde{H}$ rather than on $H$, as in Theorem (ref).

As for the EWM policy, results in Theorem (ref) can be used to derive asymptotic valid confidence intervals for $\hat{t}^s_n$ (see Appendix (ref)), and, combined with equation (ref), to characterize the asymptotic distribution of the regret $\mathcal{R}(\hat{t}^s_n)$, as derived in the next corollary.

corollaryAsymptotic distribution of regret $\mathcal{R}(\hat{t}^s_n)$ is: \begin{gather*} n \sigma_n \mathcal{R}(\hat{t}^s_n) \rightarrow^d \frac{1}{2} \frac{\alpha_2 K}{H} \chi^2\left(1,\frac{\lambda A^2}{\alpha_2 K}\right) \end{gather*} where $\chi^2\left(1,\frac{\lambda A^2}{\alpha_2 K}\right)$ is a non-centered chi-squared distribution with 1 degree of freedom and non-central parameter $\frac{\lambda A^2}{\alpha_2 K}$. The expected value of the asymptotic distribution is: \begin{align} \frac{1}{2} \frac{\alpha_2 K}{H} \left(1 + \frac{\lambda A^2}{\alpha_2 K} \right) = \frac{\alpha_2}{2} \frac{ K}{H} + \frac{1}{2} \frac{\lambda A^2}{H}. \end{align} Let $\sigma_n = (\lambda/n)^{1/(2h +1)}$ with $\lambda \in (0, \infty)$. The expectation of the asymptotic regret is minimized by setting $\lambda = \lambda^* = \frac{\alpha_2 K}{2hA^2 }$: in this case, the expectation of the asymptotic distribution scaled by $n^\frac{2h}{2h+1}$ is $ A^{\frac{2}{2h+1}} K^{\frac{2h}{2h+1}} H^{-1} C^s$, where $C^s = \frac{2h+1}{2} \left( \frac{\alpha_2}{2h} \right)^\frac{2h}{2h+1}$ is a constant not dependent on $P$.

With the optimal bandwidth $\sigma_n =O_p(n^{-\frac{1}{2h + 1}})$, the regret converges at $n^{\frac{2h}{2h + 1}}$ rate. For $h \geq 2$, this implies that the regret converges faster with the SWM than with the EWM policy: the extra smoothness assumption has been exploited to achieve a better rate for the asymptotic regret. When $h=1$, if additional regularity assumptions ensure the existence of $\tilde{H}$, the SWM policy attains the same convergence rate as the EWM policy. As for Corollary (ref), the result in Corollary (ref) does not imply convergence in the mean, and the expected value of the asymptotic distribution is reported as a measure of the location of the asymptotic distribution.

The comparison between Corollaries (ref) and (ref), which characterize the asymptotic distributions of regrets $\mathcal{R}(\hat{t}^e_n)$ and $\mathcal{R}(\hat{t}^s_n)$, highlights how the data distribution $P$ differently influences the asymptotic behavior of the regrets through $H$, $K$, which affect both distributions but with different exponents, and $A$, the bias term that affects only the SWM policy. The relevance of these asymptotic results relies on their ability to approximate behaviors in finite samples. After all, policymakers only have access to finite experimental data. Building on the theoretical results derived above, the next section uses Monte Carlo simulations to analyze the finite-sample regrets associated with the EWM and SWM policies.

Monte Carlo Simulations

I examine the finite sample properties of the EWM and SWM policies using Monte Carlo simulations. The scope of this section is twofold: first, I will provide examples of data generating processes that lead to different rankings for the two policies in terms of median asymptotic regret, to illustrate how in the asymptotic results in Corollaries (ref) and (ref) the convergence rate and the limiting distribution interact. Then, I will verify how the asymptotic results approximate the finite sample distributions, and compare the finite sample regrets of the two policies.

As data generating process, consider the following distribution $P$ of $(Y_0, Y_1, D, X)$:

align*[align* omitted — 216 chars of source]

Under $P$, the potential outcome $Y_0$ does not depend on the index $X$, and the treatment is randomly assigned with constant probability $p$. Parameter values are chosen such that $\mathbb{E}[Y_1|X=x]$ and hence $\mathbb{E}[Y_1 - Y_0|X=x]$ are increasing function of $x$, and the optimal threshold $t^*$ is 0. It can be verified that such $P$ implies the following:

align*[align* omitted — 265 chars of source]

where $\phi(t)$ and $\Phi(t)$ are the probability density function and the cumulative density function of the standard normal distribution, respectively.

I consider two models characterized by different parameter values, reported in Table (ref):

table[table omitted — 285 chars of source]

For the SWM policy, the kernel function is the cumulative distribution function of the standard normal distribution, which satisfies Assumption (ref) with $h=2$. Consequently, all analyses are conducted under Assumption (ref) with $s=h+1=3$. Asymptotic regrets are computed with the infeasible optimal bandwidth $\sigma_n^*$. Table (ref) presents the values of $K$, $H$, and $A$, along with the medians of the asymptotic regret for both policies for sample sizes $n \in \{500, 1000, 2000, 3000\}$. Compared to Model 1, Model 2 entails larger $K$ and $A$, and smaller $H$, which lead to higher median regrets under both policies. Consider the case with $n=500$. In model 1, the median of the asymptotic regret is higher with the EWM policy, whereas in model 2, it's higher with the SWM policy: this confirms that the ranking of the asymptotic median regrets depends on the unknown data distribution $P$. Because of the fastest rate, though, as $n$ increases the SWM policy exhibits relatively better performance. Regardless of the specific distribution $P$, there exists a certain sample size beyond which the asymptotic median regret with the SWM policy becomes smaller. In model 2, when $n=1,000$, the inversion of ranking already occurs.

table[table omitted — 1,130 chars of source]

To investigate the finite sample distributions of the regret, I draw samples of size $n$ from $P$ 5,000 times for each model. Each sample is used to estimate the thresholds $\hat{t}^e_n$ and $\hat{t}^s_n$. Estimating $\hat{t}^s_n$ requires specifying a bandwidth $\sigma_n$, for which I adopt the following method: I use the estimated policy $\hat{t}^e_n$ to compute $\hat{A}_{n}$ and $\hat{K}_{n}$, which are then used to compute the optimal $\hat{\lambda}_n^*$, and the optimal bandwidth $\hat{\sigma}_n^*$. In Appendix (ref), I provide and discuss formulas for estimators $\hat{A}_{n}$ and $\hat{K}_{n}$. The SWM policy $\hat{t}^s_{n}$ is hence estimated with this bandwidth $\hat{\sigma}_n^*$, which consistently estimates the optimal bandwidth $\sigma_n^*$ if $\hat{A}_{n}$ and $\hat{K}_{n}$ consistently estimate $A$ and $K$. I also consider $\hat{t}^s_{n}$ with the infeasible optimal $\sigma_n^*$ computed from the data generating process.

Estimates for $\hat{t}^e_n$ and $\hat{t}^s_n$ are used to compute regrets $\mathcal{R}(\hat{t}^e_n)$ and $\mathcal{R}(\hat{t}^s_n)$. I thus obtain the finite sample distributions of the regret, which can be compared with the asymptotic distributions derived in Corollaries (ref) and (ref). Table (ref) presents the median of these finite sample and asymptotic distributions, also depicted in Figures (ref) and (ref). Corresponding tables and figures for the mean regret are provided in Appendix (ref).

The last column of each table reports the ratio between the finite sample median regrets for the EWM and SWM policy, facilitating the comparison: a ratio larger than one indicates that the SWM policy outperforms the EWM policy. These ratios increase with the sample size, reflecting the faster asymptotic convergence rate of the SWM policy. Similar to the asymptotic results, in finite sample the SWM policy does relatively better as the sample size increases.

table[table omitted — 1,350 chars of source]
figure[figure omitted — 337 chars of source]
figure[figure omitted — 337 chars of source]

Simulations enable comparison between finite sample regrets and their asymptotic counterparts. Across all models and sample sizes, the asymptotic approximation for the feasible SWM policy (with the estimated bandwidth $\hat{\sigma}_n^*$) is relatively less accurate. Simulations suggest that this is partly attributable to the need for estimating an additional tuning parameter, the bandwidth $\sigma_n$. When the SWM policy is estimated using the infeasible optimal bandwidth $\sigma^*_n$, in fact, the asymptotic approximation is more accurate and the regret is smaller.

In Model 1, as illustrated in Figure (ref), both the finite sample and the asymptotic median regrets are lower for the SWM policy. Conversely, in Model 2 (illustrated in Figure (ref)), the finite sample median regret is lower with the EWM policy. This confirms the impossibility of ranking the policies in a pointwise sense: different distributions $P$ result in different rankings for the finite sample median regrets.

It is important to note that the ranking of the EWM and the SWM policies indicated by the asymptotic results may differ from the actual finite sample comparison. Consider, for example, Model 2 with $n=1,000$: despite the asymptotic analysis suggesting a smaller median regret with the SWM policy, the EWM guarantees a smaller regret. In this scenario, even if $P$ were known in advance, choosing according to the asymptotic approximation would not have been optimal. As $n$ increases, the approximation improves, and the rankings based on asymptotic analysis and finite sample comparisons coincide.

Monte Carlo simulations have confirmed that the asymptotic results can approximate some finite sample behavior of the regrets, highlighting some caveats to consider when applying conclusions from asymptotic analysis to finite sample regrets with the EWM and SWM policies. However, they have not yet provided insight into the practical significance of the differences between the two policies, whether these differences are relevant or negligible in real-world scenarios. An empirical illustration is useful to answer these questions, illustrating the different implications that the policies may have.

Empirical Illustration

I consider the same empirical setting as kitagawa2018should: experimental data from the National Job Training Partnership Act (JTPA) Study. bloom1997benefits describes the experiment in detail. The study randomized whether applicants would be eligible to receive a mix of training, job-search assistance, and other services provided by the JTPA for a period of 18 months. Background information on the applicants was collected before treatment assignment, alongside administrative and survey data on the applicants' earnings over the subsequent 30 months.

I consider the same sample of 9,223 observations as in kitagawa2018should. The treatment variable $D$ is a binary indicator denoting whether the individual was assigned to the program (intention-to-treat). The outcome variable $Y$ represents the total individual earnings during the 30 months following program assignment, adjusted by subtracting the average cost of the program (774 dollars) for individuals with $D=1$. This adjustment accounts for resource limitations by defining the treatment effect as positive when the benefits exceed the costs, rather than simply when the benefits are positive\footnote{My framework may not accommodate other types of budget constraints; for example, it cannot handle scenarios where only a fixed amount of resources is available for program implementation.}.

The threshold policy is implemented by considering the individual's earnings in the year preceding the assignment as the index $X$. Treatment is exclusively assigned to workers with prior earnings below the threshold, based on the expectation that program services yield a more substantial positive effect for individuals who previously experienced lower earnings. Experimental data are employed to determine the threshold beyond which the treatment, on average, harms the recipients.

To estimate the SWM policy, I use the cumulative distribution function of the standard normal distribution as the kernel function, maintaining the same assumptions as in the Monte Carlo simulations. The bandwidth is chosen using the EWM policy $\hat{t}^e_n$ to compute $\hat{A}_{n}$ and $\hat{K}_{n}$ (see Appendix (ref) for the formulas), the optimal $\hat{\lambda}_n^*$, and then the optimal bandwidth $\hat{\sigma}_n^*$. Table (ref) reports the threshold estimates, including the confidence intervals constructed as discussed in Appendix (ref). The threshold with the EWM policy is almost 500 dollars lower than with the SWM (3,107 vs 3,592 dollars), a drop of $13.5\%$. The lower threshold implies that the treatment would target fewer workers: if the EWM policy were implemented, $63.55 \%$ of the workers in the sample would receive the program services, compared to the $66.19 \%$ with the SWM policy, resulting in a $3$ percentage point difference.

table[table omitted — 580 chars of source]

Since I account for the costs of the program, the finding that the optimal threshold policy excludes certain workers from treatment does not imply that the program has a negative impact on those excluded. Rather, it likely reflects that, for workers with higher initial earnings, the program’s benefits are outweighed by its costs. However, if policymakers have different welfare objectives, they may still conclude that treating all individuals is optimal.

For these reasons, the numbers in the table should be considered with care, and clearly, the intention of this empirical illustration was not to advocate for a specific new job-training policy. Rather, the application aimed to assess if the EWM and SWM policies may have implications with relevant economic differences. Results suggest this is the case: together with theoretical and simulation findings, this implies that the choice between the EWM and the SWM policy should be thoughtfully considered, as it may determine relevant improvement in population welfare.

Conclusion

In this paper, I addressed the problem of using experimental data to estimate optimal threshold policies when the policymaker seeks to minimize the regret associated with implementing the policy in the population. I first examined the Empirical Welfare Maximizer threshold policy, deriving its asymptotic distribution, and showing how it links to the asymptotic distribution of its regret. I then introduced the Smoothed Welfare Maximizer policy, replacing the indicator function in the EWM policy with a smooth kernel function. Under the assumptions commonly made in the policy learning literature, the convergence rate for the worst-case regret of the SWM is faster than with the EWM policy. Monte Carlo simulations corroborated the asymptotic finding that the SWM policy may perform better than the commonly studied EWM policy also in finite sample. An empirical illustration displayed that the implications of the two policies can remarkably differ in real-world application.

Three sets of problems remain open for future research, to extend the findings of this paper in diverse directions. First, while my results on the rate improvement for the EWM and SWM policies are pointwise, one might ask whether they could be formalized as optimality statements. Specifically, is it possible to establish minimax rate optimality for the two policies over some family of distributions $\mathcal{P}$? For the EWM, this would relate closely to Theorems 2.3 and 2.4 in kitagawa2018should, as the assumptions underlying my $n^{\frac{2}{3}}$ rate improvement are connected to their margin assumption. For the SWM, however, the connection with the margin assumption is less clear, and alternative approaches may be considered.

Second, it would be interesting to extend the smoothing approach of the EWM policy to other policy classes. While threshold policies are convenient as they depend on a single parameter, the same intuition for smoothing the indicator function could also apply to linear index or multiple index policies, albeit with more complex derivations. Key questions include how the theory developed in this paper could be adapted to these policy classes and whether this approach might be generalized to all cases where the EWM policy is applicable. In some of these cases, the EWM policy is known to lack good computational properties. Since the SWM policy implies a smooth objective function, investigating its potential computational advantages would also be worthwhile.

Lastly, the framework developed in this paper for using experimental data to estimate optimal policies could inform experimental design. While the existing literature mainly focuses on optimal design for estimating the average treatment effect, it could be valuable to consider scenarios where estimating the threshold policy is the goal: how should the experimental design be adapted? How the allocation of units to treatment and control groups would change? The results presented in this paper, elucidating the connection between the distribution $P$ and the regret of the policy, provide a natural foundation for exploring experimental designs optimal for threshold policy estimation.