Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,239 characters · 9 sections · 57 citation commands
Statistical Treatment Rules under Social Interaction
\doublespacing
{3ex}
\setcounter{page}{1}
One of the most crucial questions for a policy maker is how to assign a treatment to an individual or a group. For example, during the COVID-19 pandemic, each government has tried to find an effective order of vaccination. Recently, statistical treatment rules based on the decision theoretic framework have received much attention in treatment evaluation studies (for a general review, see manski2004statistical,manski2021econometrics and hirano2020asymptotic). Compared to the conventional approaches based on the point estimation and inference procedures, statistical treatment rules make it possible to evaluate a broader range of treatment rules, which includes a direct map from data to an action. Despite active research in this area, most studies focus on the individualistic treatment response and we have limited results for the case where treatment outcomes depend on each other. As we can see from the vaccination example, it is important in many empirical settings to consider dependent treatment outcomes
In this paper we study a treatment assignment rule in the presence of treatment outcome dependency. In addition to the problem of vaccination, there are many applications that a policy maker has to weigh dependent treatment outcomes. \citet*{heckman1999human} evaluate the effect of a tuition reduction policy in the UK in a general equilibrium framework. They show that ignoring the outcome dependency over-estimates the effect of the policy on college enrollment more than 10 times. duflo2004scaling also argues that even a randomized control trial faces a challenge in scaling up to a larger level because of the general equilibrium effects or, more generally, dependent treatment outcomes. Using Danish data on a large job assistance program, gautier2018estimating show that the unemployed who are not selected in the program spend more time in job search than those who look for a job in provinces without such a program. Thus, the outcome of the untreated depends on that of the treated, and the treatment evaluations assuming independent treatment outcomes can mislead a policy maker.\footnote{See also beaman2012social, bursztyn2014understanding, and duflo2003role for additional examples.}
We investigate this problem in the framework of the statistical decision theory. Treatment outcomes are allowed to depend on each other in a flexible way. We aim to construct a treatment assignment rule under the minimax regret approach and to characterize it. Thus, a treatment choice using sample data, i.e.\ a statistical decision rule, is the main object of interest in this paper. Having in mind a large-scale policy implementation, we do not impose any individual network information available. Instead, we impose a shape restriction on treatment response functions following manski2013identification. Specifically, we assume anonymous interactions, which implies that the treatment response of an individual does depend on the treatment status of others but is invariant of the identity of other individuals. In other words, it is independent of the permutation of the treatment assignments on others. In the job assistance program above, for instance, this condition implies that the negative effect of the policy on the untreated only depends on the total size of people who receive the benefit of the job assistance program. This assumption provides a good approximation of the world with a large-scale policy implementation, and it makes both theoretical and empirical analyses feasible by reducing the domain of the response function substantially.
We define the sampling process carefully following the statistical decision theory framework. It contrasts to the standard individualistic treatment effect model in that our process represents both the treatment status variable and the outcome variables as a vector. The dimension of the vector is the same as the number of different treatment ratios in the target population. We adopt the minimax regret approach to handle the underlying ambiguity of the data generating process. We propose an intuitive decision rule called the multinomial empirical success (MES) rule that extends the empirical success rule in manski2004statistical to the current setup. We investigate the properties of the MES rule followed by the possible applications.
The main contributions of this paper are summarized as follows. First, we prove that the MES rule achieves the asymptotic optimality for the minimax regret criterion. Using the structure of the finite action problem in statistics literature, it extends the seminal optimality result in hirano2009asymptotics to multiple treatments. Second, we derive the non-asymptotic bounds of the expected welfare and the maximum regret under the MES rule. It is challenging to obtain these bounds since outcomes are correlated under social interaction. We also provide two applications on how these bounds can be used: (i) designing an optimal sampling procedure, and (ii) computing the sufficient sample size to allow additional covariates in the treatment rule.
The rest of the paper is organized as follows. We finish this section by reviewing related literature. In section (ref) we provide the main framework of the analysis. In section (ref) we define the MES rule and derive the upper bounds of the maximum regret. We also provide two applications of these bounds. In section (ref) we show the asymptotic optimality of the MES rule. We provide some concluding remarks in section (ref). All proofs and technical details are deferred to the appendix.
In the seminal work of manski2004statistical, he considers the statistical decision theory in the context of heterogeneous treatment rules. He proposes the empirical success rule and derives the finite sample bounds of the minimax regret. stoye2009minimax characterizes the minimax regret rule using the game theoretic approach and shows that the empirical success rule is a good approximation of the minimax regret rule under certain sampling processes. hirano2009asymptotics apply the limit experiment framework to develop large sample approximations to the statistical treatment rules.
kitagawa2018should propose the empirical welfare maximization (EWM) method that selects the treatment rule maximizing the sample analogue of the social average welfare. athey2021policy propose a doubly robust estimation procedure for the EWM problem and show the rate-optimal regret bounds. mbakop2021model consider a large class of admissible rules and propose a penalized EWM method that chooses the optimal size of the policy class. manski2016sufficient,manski2019trial argue to design clinical trials based on the goal of statistical treatment rules rather than on the statistical power of a hypothesis test. Motivated by a risk-averse policy maker, manski2007admissible and \citet*{kitagawa2022treatment} propose nonlinear transformations of welfare and regret.
manski2013identification studies identification of treatment effects with social interaction. To make the problem feasible, he proposes possible approximation methods including anonymous interaction, which will be explained in detail later. manski2009identification analyzes statistical treatment rules under the anonymous interaction assumption and the shape restriction on the mean welfare function. viviano2019policy proposes the network empirical welfare maximization method under the anonymous interaction assumption among those in the first-degree neighbor. However, our approach is different from his since it does not require heavy computation to solve an empirical optimization problem. It is also new that the proposed multinomial empirical success rule achieves the asymptotic optimality in the sense of hirano2009asymptotics.
We consider the following framework based on manski2004statistical and stoye2009minimax. Consider a social planner who assigns a binary treatment $T \in \{0,1\}$ to each individual $j$ in a heterogeneous population $J$. The population is divided into mutually exclusive and exhaustive groups based on observed characteristics (e.g.\ high school graduate vs. college graduate). Let $g \in \{1,2,\ldots,G\}$ be the index of a group and $n_{g}$ be the (population) size of group $g$. Individual $j$ in group $g$ has a response function $y_{jg} :\{0,1\}\times\{0,1\}^{n_{g} -1} \mapsto [0,1]$ that maps each possible group treatment vector $\mathbf{t} = (t_1,\dots,t_{n_{g}}) \in \{0,1\}^{n_{g}}$ into an outcome in $[0,1]$. Thus, we can write $ y_{jg}(\mathbf{t}) = y_{jg}(t_j, \mathbf{t}_{-j}) $, where $t_j$ is the treatment assigned to individual $j$ and $\mathbf{t}_{-j}$ represents the treatment vector for individuals in the same group excluding person $j$'s treatment assignment. This response function generalizes the individualistic treatment in a way that the spillover effect is allowed inside the same group (e.g.\ segmented labour markets). Note that the model allows the most flexible interactions when the whole population is categorized as a single group. The range of $[0,1]$ is a simple normalization and any bounded outcome space can be allowed. For notational simplicity, we consider a single group from now on and drop the subscript $g$ unless it causes any confusion.
We consider a probability space $(J, \Sigma, P_J)$. The population $J$ is dense in the sense that $ P_J(\{j\})=0$, for all $j\in J$. The social planner cannot distinguish members of $J$. Therefore, we can consider the model as an induced random process, $Y(\mathbf{t})$, which is a potential outcome depending not only on individual treatment status, $t_j$, but on possible treatments of other members, $\mathbf{t}_{-j}$. Given the large size of the population $J$, this random process in the most general structure is intractable. Following the social interaction literature, we impose the following assumption.
Assumption (ref) implies that a treatment ratio is a sufficient statistic for $\mathbf{t}_{-j}$. Let $\pi(\mathbf{t})$ be a treatment ratio of treatment vector $\mathbf{t}$. Then, for two treatment vectors $\mathbf{t} \neq \mathbf{t}'$ such that $\mathbf{t}= (t_j, \mathbf{t}_{-j})$ and $\mathbf{t}'= (t_j, \mathbf{t}'_{-j})$, Assumption (ref) implies that
Therefore, the outcome of a treatment $\mathbf{t}$ depends on individual's treatment status $t_j$ and $\pi(\mathbf{t})$, and we can rewrite the the response function $y_j(\mathbf{t})$ as $y_j(t_j,\pi(\mathbf{t}_{-j})): \{0,1\} \times \Pi \mapsto [0,1]$, where $\Pi:=[0,1]$. The potential outcome processes now become $(Y_0(\pi),Y_1(\pi))$ whose distribution is $P_Y(Y_0(\pi),Y_1(\pi))$. Note that the induced measure $P_Y$ can be constructed from $P_J$ given the response function $y_j(\cdot)$.
The distribution $P_Y$ is identified with a state of the world $\theta \in \Theta$ that is unknown to the policy maker. Note that $\{P_{Y,\theta}(Y_0(\pi),Y_1(\pi)): \theta \in \Theta\}$ is composed of all possible distributions on the outcome space $[0,1]^2$ for each $\pi \in \Pi$. To make the main arguments clear, we impose an additional assumption that the set $\Pi$ is discrete.
Assumption (ref) is suitable to many applied settings since the treatment ratio set may be constrained exogenously for ethical, budgetary, equity, legislative or political reasons. In addition, this is a practical assumption when experiments are costly to implement at all feasible treatment ratios. The assumption could also provide a good approximation if $\mathbf{\Pi}$ is a continuous interval but outcome function $y_j$ is smooth in $\pi$
We provide the following examples below.
We now turn our attention to a random sample that helps the policy maker infer the state of the world $\theta$. Let $\mathbf{\Pi}= \{\pi_1, \pi_2,\dots, \pi_K\}$. The experiment generates a sample space $\Omega := ( \{0,1\} \times [0, 1])^{n}$, where $n :=\sum_{k=1}^K n_{k}$ and $n_{k}$ is the subgroup size of an experiment with a treatment ratio $\pi_k$. A typical element of $\Omega$ is represented by
Conditional on the treatment $t_{n_k}(\pi_k)$, $y_{n_k}(\pi_k)$ is an independent realization of $Y_{t}(\pi_k)$ for $t=0,1$. Therefore, it helps a policy maker to infer the state of the world $\theta$. To make notation simple, we assume the equal subgroup size, $n_1=\cdots=n_K = n/K$, and $\omega^n$ is composed with n-copies of $ \omega_i := \{(t_{i}(\pi_1) ,y_{i}(\pi_1)), \dots , (t_{i}(\pi_K) ,y_{i}(\pi_K)) \}. $
The policy maker constructs a statistical treatment rule $\delta : \Omega \mapsto \mathbf{\Pi}$\, that maps a sample realization $\omega^n$ onto a treatment assignment ratio $\pi \in \Pi$. Recall that we restrict our attention to a single group in this framework but the statistical treatment rule can be group-specific when there are multiple groups. In section (ref), we extend the current frame to the multiple groups case.
The expected outcome (or social welfare) given the statistical treatment rule $\delta$ and the state $\theta$ is
where $Q_{\theta}$ is a distribution of $\omega_i$ given state $\theta$, $U(\pi,\theta):= (1-\pi) \cdot E_{\theta}[Y_0(\pi)] + \pi \cdot E_{\theta}[Y_1(\pi)]$ is the expected outcome (or social welfare) for any given treatment ratio $\pi$ in state $\theta$, and $E_{\theta}[Y_t(\pi)]$ is the mean potential outcome of treatment status $t$ given $\theta$ and $\pi$. Note that the potential outcome variable $Y_t(\pi)$ depends on the treatment of others through $\pi$. This point becomes clearer if we compare the expected outcome in (ref) with that of the individualistic treatment model (e.g.\ stoye2009minimax). When there is no social interaction, the mean potential outcome is independent of the group treatment ratio $\pi$, i.e.\ $E_{\theta}[Y_t(\pi)]=E_{\theta}[Y_t]$. Then, the expected outcome in (ref) becomes
where the last line is equal to the expected outcome in stoye2009minimax using his notation.
It is interesting to compare our framework to the individualistic multiple-treatment design. Given the finite number of treatment ratios, one might want to interpret the framework in terms of $K$ different individual treatments without any social interaction: e.g.\ define $Y_1 := Y_1(\pi_1), Y_2 := Y_1(\pi_2), \ldots, Y_K:= Y_1(\pi_K)$ and set $(Y_0,Y_1,\ldots,Y_K)$ as a vector of potential outcomes. However, this multiple-treatment design does not capture the feedback effect of the social interaction for any non-treated individual. Note that $Y_0(\pi)$ still depends on the treatment ratio $\pi$ in our framework, which is not embedded in the potential outcome vector $(Y_0,Y_1,\ldots,Y_K)$ of the standard multiple-treatment design.
The decision problem is to find a statistical treatment rule that maximizes the expected outcome function $u(\delta,\theta)$. However, there exists ambiguity in the sampling process and we need some decision criteria for unknown $\theta$. In this paper we adopt the minimax regret rule following manski2004statistical and stoye2009minimax. The regret function of $\delta$ given state $\theta$ is defined as
where $D$ is a set of all possible statistical treatment rule. The minimax regret solution of the decision problem is defined as
In this section we propose a feasible statistical decision rule and characterize it by the non-asymptotic bounds on the maximum regret. To show the main idea, we keep focusing on a single group case. The results are extended into the multiple-group case in section (ref) and we show how they can be used to determine the proper level of groups.
It is difficult to attain the optimal statistical treatment rule by solving (ref) directly since $R(\delta,\theta)$ involves integration over finite sample distributions. As an alternative, researchers may propose possible statistical treatment rules and analyze whether they achieve the optimal regret level. One of the popular rules is an empirical success rule, which substitutes empirical success rates for the population counterparts.
We propose such an empirical success rule suitable for the proposed setup. To focus on our main arguments, we restrict our attention to samples with a strict ordering of the estimates for ${U}(\pi,s)$ for all $\pi \in \mathbf{\Pi}$. We define our multinomial empirical success (MES) rule as follows:
where $\mathbf{\Pi}_{-k} := \mathbf{\Pi} \setminus \{\pi_k\}$ and
Note that, using the convention $0\cdot \infty =0$, we define $\hat{U}(0)=N_1^{-1} \sum_{n_1=1}^{N_1} Y_{n_1}(0)$ when $\pi_1=0$. Similarly, $\hat{U}(1)=N_K^{-1} \sum_{n_K=1}^{N_K} Y_{n_1}(1)$ when $\pi_K=1$.
We have a few remarks on the proposed statistical decision rule. First, we call the rule in (ref) as a Multinomial Empirical Success (MES) rule to emphasize the multinomial choice set in the setting. Second, we estimate $E_{P_{\theta}}[Y_t(\pi)]$ by using the empirical measure that depends on the unknown state $\theta$ of the world. Thus, both $\hat{U}(\pi_k)$ and the outcome of $\delta^{MES}(\cdot)$ depend on $\theta$ although it is not included as an argument explicitly. Third, the MES rule encompasses the (unconditional) empirical success rule in manski2004statistical. Let $\Pi=\{0,1\}$ with $\pi_1=0$ and $\pi_2=1$. Then, the MES rule becomes
which is the empirical success rule in manski2004statistical.
We next evaluate the expected outcome in (ref) using the MES rule in (ref):
As we discussed above, $u(\delta^{MES}, {\theta})$ is intractable since it involves all possible finite sample distributions. However, building on manski2004statistical, we can construct bounds for the expected outcome with the MES rule as follows.
It is worth comparing these bounds with those in Proposition 1 of manski2004statistical. Note that both frameworks allow the potential outcome distributions to vary across some indexing variables. For example, the potential outcomes in manski2004statistical depend on exogenous conditioning variables $X$, i.e.\ heterogeneous treatment effects over $X$. However, we focus on the dependence of the potential outcomes on the choice variable $\pi$. They look similar from the mathematical perspective, but the implications are quite different since the result in this paper allows the effect of social interaction. This point becomes clearer when we extend the model to the case that includes additional conditioning variables.
We further investigate the finite sample penalty of the lower bound in (ref), which measures the possible difference of $u(\delta^{MES}, \theta)$ from the ideal solution $U(\pi_{\pi_{M^*}}, \theta)$. First, the penalty converges to zero at the exponential rate as $N_{tk}$ increases uniformly for all $t\in \{0,1\}$ and $k \in \{1, \dots, K\}$. Second, the penalty is maximized when $\Delta_{M^*k}=\{A_{k}+A_{M^*}\} ^{1/2}/2$ for each $k \neq M^*$. Thus, we can compute the upper bound of the penalty as follows:
Third, it is interesting to investigate the relationship between the cardinality of $\Pi$ denoted by $K$ and the penalty size. Consider the following example of two possible choice sets $\Pi_1$ and $\Pi_2$ such that $\Pi_1 \subset \Pi_2$. Let $\pi_{M^*}$ be the optimal solution of $\Pi_1$. If $\pi_{M^*}$ is also the optimal solution of $\Pi_2$, then $\Pi_2$ has a larger penalty than $\Pi_1$. However, if the optimal solution of $\Pi_2$ denoted by $\pi_{M^{**}}$ is different from $\pi_{M^{*}}$, then $\Pi_2$ may have a smaller penalty than $\Pi_1$. Note that $\Delta_{M^{**}k} > \Delta_{M^{*}k}$ for all $k \in \Pi_1$ and that there may exits some $k \in \Pi_1$ such that $\exp [-2\Delta_{{M^{**}}k}^2 \{A_{k}+A_{{M^{**}}} \}^{-1} ] < \exp [-2\Delta_{{M^*}k}^2 \{A_{k}+A_{{M^*}} \}^{-1} ]$. Therefore, a larger choice set may improve the finite sample lower bound only if it contains a better welfare outcome. Finally, we investigate the uniform bound of the regret function over $\theta$. The upper bounds of the regret function with $\delta^{MES}$ is represented in terms of the penalty:
Different from the result in manski2004statistical, $A_{{M^*}}$ in the right hand side depends on $\theta$ since $\pi_{M^*}$ is defined in terms of $U(\pi,{\theta})$. Therefore, we need an additional step to achieve the uniform bound. Let $\overline{A} := \max_{k \in \{1,\dots, K\}}A_k$. Note that $\overline{A} \ge A_{M^*}$ and that $\overline{A}$ is independent of $\theta$. Then, the desired uniform bound is achieved as follows:
These finite sample bounds give us two useful applications. First, we apply this bound to solve the quasi-optimal experiment design problem. Second, we can extend the bound to the covariate dependent treatment rule and determine the minimum sample size to adopt a finer covariate set as in manski2004statistical. We provide these applications in the following two subsections.
We study the optimal experiment design problem under interference using the upper bound of the maximum regret. Specifically, we focus on the randomized saturation design which is composed of two-stage randomized experiments (for example, see baird2018optimal). Suppose that we are given many clusters. In the first stage, we assign different treatment ratios in $\Pi$ to each cluster randomly according to a probability distribution $f$. In the second stage, a binary treatment is assigned to each member of a cluster according to a treatment ratio assigned in the previous stage. Therefore, the randomized saturation design is fully characterized by a pair $(\Pi,f)$ and it encompasses other designs like clustered, block, and partial population designs commonly employed under interference.
We now consider an experiment design problem that minimizes the maximum regret. We cannot compute the exact regret function because of the ambiguity in $\theta$. Instead, we reformulate the problem as minimizing the feasible upper bound of the regret in (ref).
Recall that $N$ denotes the total sample size over all clusters and $\Pi=\{\pi_1, \pi_2, \cdots, \pi_K\}$ be a finite set of treatment ratios. Since $\Pi$ is a finite set, we can write $f=\{(\alpha_1, \alpha_2, \cdots \alpha_K): \sum_{k=1}^K \alpha_k = 1\}$, where $\alpha_k$ is a probability mass of assigning $\pi_k$. The subsample sizes can be written in terms of the treatment ratios and their corresponding probabilities: $N_{k0}= (1-\pi_k)\alpha_k N$ and $N_{k1}= \pi_k\alpha_kN$ for all $k= 1, 2,\ldots, K.$ Then, for each $A_k$, we have
which makes the optimization problem simple. Without loss of generality, let $\overline{A}=A_1$. We substitute $A_k$ in (ref) and drop all irrelevant variables to get
Solving this optimization problem, we derive the quasi-optimal design of equal $\alpha_k^*$ ($\alpha^*_k=1/K$) only when $K=2$. It is worthwhile to note that baird2018optimal derive the optimal randomized saturation design based on the statistical power but we focus on the maximum regret directly (see manski2016sufficient for further discussion).
In this section, we extend the model and consider covariate-dependent treatment rules. We first introduce new notation. Let $X$ be a vector of covariates. In the similar spirit of Assumption (ref), we restrict our attention to discrete and finite covariates. Then, we can vectorize the possible outcomes of $X$ and partition the population into $L$ different subsets denoted by $\mathcal{X}:=\{x_1,\ldots,x_L\}$. To make notation simple, we assume a common domain of treatment ratios $\Pi$ for each $x_l$\footnote{We can allow different assignment ratio sets at the cost of extra notation, e.g.\ $\Pi:=\cup_{l=1}^L \Pi_l$, where $\Pi_l:=\{\pi_1,\ldots,\pi_{K_l}\}$ is the set of assignment ratios for $x_l$.}. We define a statistical treatment rule as $\delta(x,\omega^n):\mathcal{X}\times \Omega \mapsto \Pi$. Let $\boldsymbol{\pi}:=(\pi_1,\ldots,\pi_L)'$ be a vector of treatment assignment ratios, where $\pi_l$ is applied to subgroup $x_l \in \mathcal{X}$. Let $\boldsymbol{p}$ be a vector of population subgroup proportions. Then, $\bar{\pi}:= \boldsymbol{p}'\boldsymbol{\pi}$ becomes the unconditional treatment ratio. Under Assumption (ref), the response function can be rewritten as $y_j(t_j, \bar{\pi})$.
Given $\boldsymbol{\pi}$ and $\theta$, the outcome of the subgroup with covariate $x_l$ is
Note that $U_l$ is affected by the treatment ratios of other covariate types through $\bar{\pi}$ as well as its own ratio $\pi_l$. Let $\boldsymbol{\delta}(\omega^n):=(\delta(x_1,\omega^n), \ldots, \delta(x_L,\omega^n) )$ be a vector of statistical treatment rules over $\mathcal{X}$ when sample $\omega^n$ is realized, i.e.\ $\boldsymbol{\delta}(\omega^n): \Omega \mapsto \Pi^L$. The expected outcome of the whole population is defined by the weighted sum of $U_l$:
If $\Pr(X=x_l;\theta)=1$ for some $l$, then $\pi_l=\bar{\pi}\equiv \pi$, $L=1$ and $u(\delta,\theta)=\int U(\delta(\omega^n),\theta)dQ_{\theta}^n$. Therefore, the expected outcome becomes equation (ref), where there exists a single type of population.
Similar to (ref), we can define the minimax regret solution of the decision problem as
where $R(\boldsymbol{\delta}, \theta):= \max_{ \boldsymbol{d} \in \boldsymbol{D}} u(\boldsymbol{d},\theta) - u(\boldsymbol{\delta},\theta)$ is a regret function. Since the expected welfare with covariate $x_l$ is affected by the treatment assignment ratios of other covariates $x_m\neq x_l$, we need to find the decision rule simultaneously over all elements in $\mathcal{X}$, i.e.\ the decision rule vector $\boldsymbol{\delta}$. It is worth noting that, when we consider $x_l$ as a single group, this extension can be interpreted as multiple groups with interaction between groups via $\bar{\pi}$.
We now construct the multinomial empirical success rule conditional on covariate $x_l$. Note that $\Pi^L$ contains at most $K^L$ elements, $\vert \Pi^L \vert = K^L < \infty$. Let $\boldsymbol{\pi}_k$ be a generic element of $\Pi^L$. Then, the population (unconditional) treatment ratio is $\bar{\pi}_k=\boldsymbol{p}'\boldsymbol{\pi}_k$. The empirical mean of $Y_t(\bar{\pi}_k)$ conditional on $x_l$ is
Finally, the conditional multinomial empirical success rule (CMES) is defined as follows:
where $\Pi^L_{-k} := \Pi^L \setminus \{\boldsymbol{\pi}_k\}$ and
where $\pi_{kl}$ is the $l$-th element of the $L$-dimensional vector $\boldsymbol{\pi}_k$. The CMES rule $\delta^{CMES}(\omega)$ in (ref) looks similar to the (unconditional) MES rule in Section (ref). However, $\boldsymbol{\pi}_k$ is now an $L$-dimensional vector and the rule itself is an $L$-dimensional vector-valued function. Let $U(\boldsymbol{\pi}_k, \theta)$ be the population counterpart of $\hat{U}(\boldsymbol{\pi})$ by replacing $\hat{E}_{\theta}$ with $E_{\theta}$. Then, we can define the expected outcome given the CMES rule $\boldsymbol{\delta}^{CMES}$ as follows:
We are now ready to extend the the bounds of the expected outcome in ((ref)) to the CMES rule.
Using the similar arguments in Section (ref), we define the non-negative finite sample penalty:
and derive the following inequality:
Then, we can derive the uniform bound of the regret function, which can be recovered from the observable:
where $\bar{A}_l:= \max_{k \in \{1,\dots, K^L \}}A_{kl}$ $ \forall l \in \mathcal{L}.$
We next investigate the relationship between the sample size and the proper conditioning level of covariates. Recall that given a fixed sample size using all available covariates may reduce the statistical precision in practice. Let $\mathcal{Z} := \{z_1,\dots z_{L'}\}$ be a partitioning of the covariate space that is coarser than $\mathcal{X}$. Thus $L'< L$ and there exists a mapping $z(\cdot) :\mathcal{X} \mapsto \mathcal{Z}.$ Slightly abusing notation, we use the same $\boldsymbol{\pi}$ and $\boldsymbol{p}$ for assignment ratios and proportions whose dimension is $L'$. Finally, if $\boldsymbol{\pi}_{k'}$ is a generic element of $\Pi^{L'}$ and $\boldsymbol{\delta}_Z^{CMES}$ is the MES rule conditional on $Z$, then the population expected outcome becomes:
where $ U(\boldsymbol{\pi}_{k'}, \theta) := \sum_{l'=1}^{L'} \Pr(Z=z_{_{l'}})\cdot U_{l'}(\boldsymbol{\pi}_{k'},\theta) \equiv \sum_{l'=1}^{L'} \Pr(Z=z_{_{l'}})\cdot[(1 - \pi_{k'l'}) \cdot E_{P_{\theta}}[Y_0(\bar{\pi}_{k'})|Z=z_{_{l'}}] + \pi_{k'l'} \cdot E_{P_{\theta}}[Y_1(\bar{\pi}_{k'})|Z=z_{_{l'}}] ]$ and $\bar{\pi}_{k'}:= \boldsymbol{p}'\boldsymbol{\pi}_{k'} ,$ $ k' \in \{1, \dots, K^{L'}\} $ and $ l' \in \{1, \dots, L'\}.$
Similar to the results in Theorem (ref), we can bound the expected outcome in the following corollary.
We now suppose that the decision maker need to choose the conditioning level between $X$ and $Z$. The idealized bounds for the regret function is as follows.
where $$D(\theta):=\sum_{k'=1}^{K'} \exp \left( -2\Delta_{M^{**}k'}^2\cdot\left\{\sum_{l'=1}^{L'} \Pr(Z=z_{_{l'}})^2 (A_{k'l'}+A_{M^{**}l'})\right\}^{-1} \right)\cdot \Delta_{M^{**}k'}.$$ Finally, we achieve a uniform bounds on the maximum regret function as follow.
where $$L:=\sup_{\theta \in \Theta} \left\{\sum_{l=1}^L \Pr(X=x_l)\cdot U_l(\boldsymbol{\pi}_{M^{*}}, \theta)-\sum_{l'=1}^{L'} \Pr(Z=z_{_{l'}})\cdot U_{l'}(\boldsymbol{\pi}_{M^{**}}, \theta)\right\},$$ and $$H:=\sup_{\theta \in \Theta} \left\{\sum_{l=1}^L \Pr(X=x_l)\cdot U_l(\boldsymbol{\pi}_{M^{*}}, \theta)-\sum_{l'=1}^{L'} \Pr(Z=z_{_{l'}})\cdot U_{l'}(\boldsymbol{\pi}_{M^{**}}, \theta) + D(\theta)\right\}.$$ Using these bounds, we can compute the minimum sample size to test the proper level of conditioning variables. Let $\boldsymbol{N}_{KTL}:= \left(N_{ktl}: k=1,\ldots,K, t=0,1,\mbox{ and } l=1,\ldots,L\right)$ be a 3-dimensional array of stratum sample sizes. Recall that the upper bound of the maximum regret conditional on $X$ decreases as each $N_{ktl}$ increases. Therefore, we can find a sufficient sample size that justifies conditioning on $X$ rather than conditioning on $Z$:
where we minimize each component of vector $\bold{N}_{TKL}$. Similar to the results in manski2004statistical, it requires additional bound conditions on $U_l(\pi_{M^*},\theta)$ and $U_{l'}(\pi_{M^{**}},\theta)$ to solve for $\boldsymbol{N}_{KTL}$. Note also that the solution may not be unique since $\boldsymbol{N}_{KTL}$ is a tensor.
In this subsection, we conduct some numerical experiments, where we determine a sufficient sample size to use covariate-dependent treatment rules. Suppose that we have a binary covariate $X=\{low,high\}$ available in a sample. We now construct a treatment rule with or without the covariate. The sufficient sample size guarantees that the maximum regret from a covariate-dependent rule is smaller than that from a rule without considering any covariate. Thus, we can focus on covariate-dependent rules if the sample size is bigger than the sufficient one.
In this experiment, a sample is partitioned into 2 groups ($X=low$, $X=high$), and $L = \vert \mathcal{X} \vert= 2$. Therefore, covariate-dependent rules $\boldsymbol{\pi}$ also becomes a 2-dimensional vector $\boldsymbol{\pi}=(\pi_{low},\pi_{high})$. Suppose that we consider two possible treatment rules, $\{\boldsymbol{\pi}_1= (0.5, 0.5), \boldsymbol{\pi}_2= (0.7, 0.3) \}.$ Unconditional treatment ratios for them becomes:
We set that $\Pr(X=low)$ varies in $\{0.1, 0.5, 0.9, 0.99\}$ and that $\Pr(X=high):=1-\Pr(X=low)$ varies in $\{0.9, 0.5, 0.1, 0.01\}$. Recall that $(N_{ktx}: k=1,2, t=0,1, \mbox{ and } x= low, high)$ denotes the sample size of each partition separated by treatment rule $k$, treatment status $t$, and covariate $x$. In addition, $N$ denote the total sample size. $N_1$ and $N_2$ denote the sample sizes of each cluster, where we apply $\boldsymbol{\pi_1}$ and $\boldsymbol{\pi_2}$, respectively. Assuming that all states of the nature are feasible, we compute the lower bound of maximum regret for the MES rule that does not depend on covariate $X$. We also compute the upper bounds of maximum regret for the covariate-dependent MES rule as the sample size increases. We then check when this upper bound with covariates becomes smaller than the lower bound without covariates.
In Tables (ref)--(ref), we summarize the experiment results. We denote the upper bound with $X$ in bold when it becomes smaller than the lower bound without $X$. The sufficient sample size is as low as $N=21$ when $\Pr(X=low)= 0.1$, $N=18$ when $\Pr(X=low)= 0.5,$ $N=68$ when $\Pr(X=low)= 0.9,$ and $N=5875$ when $\Pr(X=low)= 0.99.$ In each table, We also provide a breakdown of the sample sizes in each partition. This numerical study shows that covariate-dependent treatment rules can be justified with relatively small sample sizes unless the sizes of heterogeneous groups are quite uneven, e.g.\ $\Pr(X=low)=0.99$.
In this section, we study the asymptotic optimality of the multinomial empirical success (MES) rule. We first transform the multivariate decision problem into a vector-valued binary decision problem. Then, we show the asymptotic optimality of MES by extending the limit experiment framework in hirano2009asymptotics into the vector-valued binary decision problem.
We first define $K(K-1)/2$-dimensional vector
where $\delta_{n,(k,k')}= \mathbbm{1}(\hat{U}(\pi_k)>\hat{U}(\pi_{k'}))$. We will call $\delta^{VMES}_n$ the vectorized multinomial empirical success (VMES) rule.\footnote{We use the subscript $n$ hereafter to distinguish a finite sample decision rule from the corresponding asymptotic one.} Note that the VMES rule has $2^{K (K-1)/2}$ different actions while the MES rule has only $K$ actions. However, a set of actions by the VMES rule is always uniquely mapped into an action by the MES rule since it gives us the preference order among all $K$ actions. To the best of our knowledge, this is the first paper to investigate the asymptotic optimality of a multiple statistical decision problem by transforming it into a vector-valued binary decision problem.\footnote{A similar idea has been in the multiple hypothesis testing literature for a long time, where they convert a $K$-multiple hypothesis problem into a $2^K$-finite action problem (see, e.g.\ lehmann1952testing,lehmann1957theory and cohen2005decision). }
We now have $J:=K (K-1) /2$ binary decision problems. Following van1991asymptotic and hirano2009asymptotics, we investigate the asymptotic optimality around the local alternatives. We first restrict our attention the parametric class of $Q$ whose extension to the semiparametric class follows immediately. Let $\mathcal{E}_n:=\{ Q^n_{{\theta}}:{\theta} \in \Theta \subset \mathbb{R}^d \}$ be a sequence of experiments, where $\Theta$ be an open subset of $\mathbb{R}^d$. We define a vector of welfare contrasts
where $g_j(\theta):=U(\pi_k,{\theta}) - U(\pi_{k'},{\theta})$ is the welfare contrast between $\pi_k$ and $\pi_{k'}$. For notational simplicity, we use $j$ for generic combination $(k,k')$, where $j=1,\ldots, J$ and $J=K(K-1)/2$. We consider local alternatives around $\theta_0$, where $g({\theta}_0) = 0$. This local problem is the most difficult case in the parameter space. If $g_j({\theta}_0)\neq 0$ for a given $\theta_0$, one action is strictly dominated by the other around $\theta_0$ and the decision between $(k,k')$ becomes trivial asymptotically.
We next define a loss function. We consider a loss function that is additively separable for each binary decision problem $j$:
where $L_j$ is a loss function for a binary decision rule $\delta_j$ between $\pi_k$ and $\pi_{k'}$. Specifically, we use the regret loss in this analysis:
Using the loss function in (ref) and experiment $Q_{\theta}^n$, we define a risk function as usual:
Note that the risk function is also additively separable.
To achieve a tractable asymptotic experiment, we assume that $Q_{\theta}$ is differentiable in quadratic mean (DQM) at $\theta_0$. For the formal definition, let $q_{\theta}$ be the density of $Q_{\theta}$ with respect to Lebesque measure $\mu$. Then, there exists a measurable function $s(\omega)$ such that, as $h \to 0$,
We can usually compute $s(\omega)$ by $s = \frac{\partial \log q_{\theta}}{\partial \theta}\vert_{\theta=\theta_0}$, and the Fisher information matrix is defined as $I_0=E_{\theta_0}[ss']$. Applying the standard local asymptotic normality arguments, we can show that the limit experiment becomes $N(\Delta|h,I_0^{-1})$, i.e. the multivariate normal distribution with mean $h$ and variance $I_{0}^{-1}$ (see Proposition 3.1 in hirano2009asymptotics).
We next define the corresponding loss and risk functions in the limit experiment. Recall that $g(\theta)$ is a $J\times 1$ vector of welfare contrasts with $g(\theta_0)=0$. Let $\triangledown_\theta g$ be a $J\times d$ matrix of partial derivatives of $g$ at $\theta_0$. Then, under some smoothness assumption on $g$, we have $\sqrt{n}g(\theta_0+h/\sqrt{n}) \to (\triangledown_{\theta} g)h$. Then, we observe that
where $\triangledown_{\theta} g_j$ is the $j$-th row of matrix $\triangledown_{\theta} g$. Using the additive separability, we can define the asymptotic loss function as
Similarly, we can define the corresponding asymptotic risk function as
Abusing notation slightly, we use the same $\delta_j$ for both $R_{j}(\delta_j,\theta)$ and $R_{j,\infty}(\delta_j,h)$. However, notice that one in $R_j$ is defined on the sample sample $\omega^n \in \Omega$ while the other in $L_{j,\infty}$ is on the simpler asymptotic experiment space $\Delta \in \mathbb{R}^d$.
In the next theorem we characterize the functional minimization problem in the limiting Gaussian experiment. We first define additional notation. Let $h_0$ be a vector such that $(\triangledown_{\theta} g) h_0 = 0$. For any $b_j \in \mathbb{R}$, we slice the parameter space as follows
Note that parameter $h_j(b_j,h_0)$ in the slice satisfies $(\triangledown_{\theta} g_j)h_j = b_j$, which is the $j$-th component of the welfare contrast vector $g(\theta)$. Because of the additive risk function and the Neyman-Pearson lemma, we can characterize the minimization problem by investigating a set of threshold rules over a vector of the sliced parameter space, separately.
Condition (ref) requires that higher loss be assigned to any incorrect choice for each $j$, and $L_{j,\infty}(\delta_j,h)$ clearly satisfies the condition. Since the expectation is a linear operator the additively separable loss function in (ref) assures the additive separability of risk function $R_{\infty}(\delta,h)$. Theorem (ref) (i) implies that threshold rule $\delta_{c}(\Delta)$ is admissible on the subspace. This result is an extension of the Theorem 3.4 in hirano2009asymptotics into a finite action problem.
To finalize our arguments on the minimax optimality, we collect all the regularity conditions.
These regularity conditions are similar to those in hirano2009asymptotics. Assumption (ref) is a mild extension of the welfare contrast to a vector-valued function. We impose that the smoothness assumption holds element-by-element. Assumption (ref) is the standard condition for the local asymptotic normality. Therefore, the asymptotic experiment can be approximated by the multivariate normal distribution for each $j$. Finally, Assumption (ref) assures the existence of an efficient estimator for $\theta_0$ and a consistent estimator for $\sigma_{g_j}$ for each $j$.
This theorem is a gentle extension of Theorem 3.5 of hirano2009asymptotics to the finite action framework with the additively separable loss function. In this paper, we focus on the statistical decision problem under social interaction, where it is transformed into choosing the fraction of the treatment. However, the result of this theorem is applicable to any case, where the decision problem is represented as a choice among multiple actions.
Straightforward is an extension to semiparametric models. We have restricted our attention to the class of parametric models $\Theta$ in this section, but we can extend it to the class of distributions $\mathcal{P}$ with more complicated notation. Instead of repeating the same arguments with messier notation, we refer to hirano2009asymptotics and van1991asymptotic for the extension. The main difference is that the multivariate Gaussian limit experiment is now replaced by an infinite Gaussian sequence.
Since the optimal decision rule has the same threshold constant both in parametric models and semiparametric models, we can claim the asymptotic optimality of the MES rule based on the finite action framework. Suppose that we have a random sample $(y_t(\pi_k), y_t(\pi_{k'}))$ for the binary decision problem between $\pi_k$ and $\pi_{k'}$. Let $F_{t,k}$ and $F_{t,k'}$ be the distributions of the sample, which is unknown but included in the class of $\mathcal{P}$. We assume that $\mathcal{P}$ is the largest class of distributions satisfying
Recall that the welfare contrast function in this binary decision problem becomes
Note that the MES rule can be written as
where
Since $\hat{g}_{n,j}$ is an asymptotically efficient estimator of $g_j$ bickel1993efficient, we can conclude that $\delta_{n,(k,k')}$ is asymptotically minimax optimal for the regret loss function and that the MES rule is asymptotically optimal for the additively separable loss function.
In this paper we study statistical treatment rules under social interaction. We impose the anonymous interaction assumption, and consider a treatment decision problem, where we choose the treatment ratio for each cluster. We propose a simple but intuitive rule called the multinomial empirical success (MES) rule. We construct the finite sample regret bound of the MES rule and show how it can be applied in the treatment decision problems. Finally, we show that the proposed MES rule achieves the asymptotic optimality in the sense of hirano2009asymptotics.
We may consider a few possible extensions. It is interesting to investigate the finite sample optimality of the MES rule. It does not work immediately if we apply the finite action problem framework, which we adopt in the asymptotic optimality analysis, and the game theoretic approach in stoye2009minimax in the finite sample case. It is also interesting to relax the anonymous interaction assumption. Then, we have to ask what kind of additional information help reduce the dimension of the action space. The network information can be such an example. We leave these questions for our future research.