Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,938 characters · 11 sections · 48 citation commands
Incorporating Preferences Into Treatment Assignment Problems
Preferences for treatments often affect their efficacy. Being assigned a disliked treatment makes an individual less motivated and less tolerant of any difficulties or inconveniences involved in that treatment. Cook1979 termed such a phenomenon “resentful demoralization.” The presence of resentful demoralization possibly leads to the heterogeneity of the treatment effect with respect to the preference. This heterogeneity is sometimes called the preference effect. The existence of the preference effect has been observed in several fields, including education Wing2017,Little2008, medical care Long2008,Zoellner2019, and energy saving programs Ida2022. In the presence of the preference effect, individualizing the treatment assignment for each true preference type is effective. If two treatments exist, say treatment 1 and 0, an example of the individualized assignments is the one giving treatment 0 to individuals preferring treatment 1 and giving treatment 1 to individuals preferring treatment 0. Such individualized treatment assignments based on individual characteristics (i.e., covariates) are called individualized treatment rule (ITR). Given any welfare function (typically, population mean outcome), the goal of individualization is to maximize welfare. The data-dependent decision of the ITR has been studied in the growing literature on statistical treatment choice Manski2004,Athey2021,Kitagawa2018,Mbakop2021,Hirano2009. The literature usually assumes that the covariates used for individualization are observable when an ITR is to be implemented. In the current study's context, this means that the true preference type is assumed to be observable. However, the true preference is private information and unobservable in nature. Instead, to implement the ITR, we must rely on the stated preference by asking individuals. The true and stated preferences are not necessarily the same. On the contrary, when individuals are informed in advance about the ITR, some individuals have a strong incentive to tell a lie. For instance, recall the example of an ITR in the previous paragraph, that is, the ITR that gives the converse treatment to the preferred one. Consider individuals who prefer treatment 1 and suppose they know the ITR and are asked about their preference. Their truthful preference revelation gives them treatment 0, while the false preference revelation gives them the preferred treatment. Thus, telling a lie becomes the optimal behavior for these individuals. Recently, several studies have analyzed individualized assignment problems with strategic agents Munro2023,Sahoo2022,Harris2023. In these problems, individuals have knowledge of the incoming ITR and strategically choose the values of their own covariates (not necessarily stated preference). Viewing stated preferences as a covariate, we can interpret our assignment problem (i.e., individualized assignment using stated preferences) as an instance of individualized assignment problems with strategic agents. Unfortunately, existing studies are not relevant to the current problem. This is because those studies have implicitly or explicitly assumed heterogeneous costs for choosing the difference level of covariate values. This presumption is adequate if ITRs use covariates such as test scores because improving test scores usually requires study effort. However, this presumption is inappropriate in the current problem as the preference statement is costless. Nevertheless, many real-world examples where the assignment of objects is determined based on the stated preference can be observed. A leading example is assignments of public schools to students Abdulkadiroglu2005,Abdulkadiroglu2005a. These assignments are based on the stated preference for schools. This study investigates the treatment assignment problem where ITRs use stated preferences and individuals know the applied ITR before the preference statement. First, we formally model the treatment assignment problem to individuals with preferences for treatments, which we briefly describe. There exist two treatments, treatment 1 and treatment 0, and each individual has a strict preference for treatments. The strictness excludes the indifference between distinct treatments. Hence, there exist two types of individuals in terms of preferences: individuals who strictly prefer treatment 1 and those who strictly prefer treatment 0.\footnote{The strictness of preferences is a common presumption in the literature of matching markets Gale1962,Ergin2002,Roth1982.} In this case, an ITR is described by the probability of giving treatment 1 for each preference type.\footnote{For simplicity, we do not consider covariates other than preference. The results of this study can accommodate covariates other than preferences as long as the additional covariates are discrete and not manipulatable.} In other words, an ITR is a pair of lotteries over treatments; one is given to individuals preferring treatment 1, and the other is given to individuals preferring treatment 0. The welfare function is set to the population mean outcome, following standard practice. The critical assumption on individuals’ preference statement is that each individual prefers the lottery that gives the preferred treatment with a higher probability. Equivalently, each individual maximizes their own expected utility, a standard assumption in microeconomics. The first result, (ref), gives the optimal ITR that maximizes welfare when individuals respond strategically to ITRs. The result leads to three findings. First, the knowledge of the \emph{true} preference type distribution and the conditional average treatment effect (CATE) given each \emph{true} preference type suffices for constructing the optimal ITR. The optimal ITR is determined by the signs of the CATEs multiplied by the share of the corresponding preference type. Second, the oracle ITR differs from the naive ITR that maximizes welfare while ignoring individuals' strategic preference statements. This suggests the significance of individuals' strategic behavior. Last, the optimal ITR is \emph{strategy-proof}, that is, no individual has a strong incentive to make a false preference statement. The strategy-proofness is regarded as a desirable property because strategy-proof ITRs reduce the burden of individuals’ thoughts. No matter how individuals contrive a scheme, there is nothing more to gain than to express their true preference. To construct the optimal ITR, the distribution of the true preference type and the CATEs given the true preference type are necessary. This information is unknown in practice and must be identified and estimated from data. Unfortunately, however, the identification is not straightforward. For example, data in which individuals freely choose the preferred treatment is useless because no individual experiences the converse treatment to the preferred one. A naive idea that seems to work is to conduct a randomized controlled trial (RCT) with a pre-treatment survey on preference Torgerson1996. In this RCT, the pre-treatment survey first asks for the preferred treatment. Then, conditional on the answered preference type, treatments are randomly assigned. Unfortunately, this experiment does not necessarily identify the objects of interest, for the survey responses are stated preferences. The stated preferences do not necessarily correspond to true preferences unless the true preference revelation is adequately incentivized. For example, Torgerson1996 conducted an RCT with the pre-treatment survey, where participants were informed that the treatments were randomly assigned with equal probability and asked about their treatment preferences. As the assignment probability did not depend on the stated preference, any preference statement was optimal behavior. As a result, the participants might have stated their preferences falsely. To overcome the difficulty above, we introduce two particular experimental designs that allow us to identify the true preference type distribution and the CATE given the true preference type. One is the \emph{strictly strategy-proof RCT (SSP-RCT)}, an adjustment of the RCT with a pre-treatment survey so that the true preference revelation becomes the strictly optimal behavior for any individual. We demonstrate that under the assumption of individuals’ expected utility maximization, the SSP-RCTs can identify the objects of interest. The other experimental design is the \emph{doubly randomized preference trial (DRPT)} Janevic2003,Rucker1989,Wennberg1993. The DRPT randomly assigns individuals to treatment 0, treatment 1, and free-choice groups; in the former two groups, the treatment exposure is exogenously determined, while individuals’ choice determines it in the third group. As illustrated in Ida2022,Wing2017,Long2008, the DRPTs---combined with the exclusion restriction Angrist1996---can identify the target parameters. Presuming data derived from data generating processes like the SSP-RCTs or DRPTs, we develop data-dependent procedures to determine an ITR, that is, \emph{statistical treatment rules (STRs)}. Following Manski2004, we evaluate the statistical performance of our proposed STRs with the maximum regret, the worst-case loss incurred due to not knowing the true data generating process. Specifically, as in Kitagawa2018, we derive finite-sample upper bounds of the maximum regret. These results imply that the maximum regret of the proposed STRs converges to zero at a rate of the square root of the sample size. \paragraph{Related Literature}{This study contributes to the literature on individualized treatment assignment problems with strategic agents Munro2023,Sahoo2022,Harris2023. As mentioned above, the results of these studies do not apply to the problem this study addresses. This is because these studies focus on covariates that require some cost for manipulation, while the preference statement can be made without any cost. Moreover, the model has other differences. Harris2023 consider a dynamic model while the deployed ITR is fixed over time. They assume that individuals have a homogeneous preference for treatments. Sahoo2022 consider a dynamic model where the implemented ITR is consecutively updated and assumes a homogeneous treatment preference. Their model also incorporates capacity constraints to capture individuals' competition for scarce treatment. Unlike those two studies, this study develops a static model in which individuals have heterogeneous treatment preferences and no capacity constraints exist. The model of this study is very similar to that of Munro2023, the only difference being the cost of manipulating the covariates. See (ref) for details. This study is also related to experimental designs incorporating individuals' preferences. Brewin1989 propose an experimental design in which only individuals who are indifferent between treatments are randomly assigned; the other individuals are given the preferred treatment. Zelen1990 proposes randomized consent designs where individuals are allowed not to comply with the randomly assigned treatment. The experimental designs proposed in Rucker1989,Wennberg1993,Janevic2003 can be classified as DRPTs; that of Wennberg1993 is most similar to the DRPT in this study. Torgerson1996,Torgerson1998 discuss a combination of conventional RCTs with a pre-treatment survey on preferences. However, they do not consider the possibility that true and stated preferences may differ and thus do not discuss how to make individuals express true preferences. Narita2021 also proposes a variant of RCT with a pre-treatment survey that achieves the Pareto efficiency among participants. However, the design is not totally strategy-proof but only approximately strategy-proof. The SSP-RCTs proposed in this study are versions of RCT with a survey that incentivizes truthful preference revelation at the expense of Pareto efficiency. } \paragraph{Organization of Paper}{The rest of this paper is organized as follows. (ref) formally models the individualized treatment assignment problem incorporating treatment preferences. We derive the optimal ITR that maximizes welfare under individuals' strategic preference revelation. (ref) discusses two experimental designs---the SSP-RCTs and DRPTs---that allow us to identify the distribution of true preference type and conditional average treatment effect given the true preference. Then, we propose the STRs associated with data from the SSP-RCTs and DRPTs. In addition, we evaluate the statistical performance of the proposed STRs with the maximum regret. (ref) demonstrates the usefulness of the proposed STR using the results reported in Wing2017. (ref) concludes this paper. All proofs are relegated to (ref). }
This section models the individualized treatment assignment problem with treatment preferences and derives the optimal ITR. Specifically, (ref) develops the model and (ref) discusses the optimal ITR that maximizes welfare.
We first discuss the standard model that assumes the observability of true preference. Then, we modify the model for the case when true preferences are not observable but stated preferences are observable.
Suppose two treatments exist, elements of $\mathcal{D}=\{0,1\}$. A policymaker plans to assign one of the treatments to each individual in the population of interest. The population is modeled as a probability space $(I,\Sigma,\mathbb{P})$, where $I$ denotes the set of individuals. For each treatment $d \in \mathcal{D}$, each $i \in I$ has a potential outcome $Y_i(d) \in \mathbb{R}$ that would be realized if $i$ was assigned treatment $d$. We maintain the stable unit treatment value assumption throughout the study. Suppose each individual has a strict preference $\succsim_i$ for treatments (i.e., complete, transitive, and antisymmetric binary relation defined over $\mathcal{D}$). For treatments $d$ and $d'$, $d \succsim_i d'$ denotes that $i$ prefers $d$ to $d'$. The antisymmetric part of $\succsim_i$ is denoted by $\succ_i$. Note that the antisymmetricity rules out indifference between distinct treatments. Hence, there are only two types of individuals in terms of preferences: type 1 is those who strictly prefer 1, and type 0 is the converse. We denote the true preference type of individual $i$ by $T_i \in \mathcal{T} \coloneqq \{1,0\}$. The policymaker does not know the tuple $(Y_i(0), Y_i(1), T_i)$ for any individual $i$. Instead, suppose that the policymaker knows the joint distribution of $(Y_i(0), Y_i(1), T_i)$. Based on this information, the policymaker determines the probability of giving treatment 1 for each preference type. Formally, the policymaker chooses an individualized treatment rule (ITR), $\delta:\mathcal{T} \to [0,1]$, where $\delta(t)$ is the probability of giving treatment $t$ to individuals with preference type $t$. The standard treatment assignment problem Manski2004, Kitagawa2018, Athey2021 assumes that all of the pre-treatment individual characteristics used for an ITR are observable when the ITR is to be implemented. In the current setup, this means that the true preference type, $T_i$, is observable for any individual. Then, given an ITR $\delta$, the individual $i$'s treatment is drawn from the Bernoulli distribution with parameter $\delta(T_i)$. We refer to this setup as the environment under true preference observation. The policymaker desires an ITR that maximizes the welfare defined as the expected outcome attained under an ITR. In the environment under true preference observation, given a joint distribution of $(Y_i(0), Y_i(1), T_i)$, the welfare under an ITR $\delta$ is
The law of iterated expectation yields
where $\tau(t) \coloneqq \mathbb{E}[Y_i(1) - Y_i(0) | T_i = t]$ denotes the conditional average treatment effect (CATE) for true preference type $t$. From (ref), we can easily describe the ITR that maximizes the welfare $W_{\operatorname{T}}$. Specifically, an ITR $\delta$ maximizes $W_{\operatorname{T}}$ if and only if it takes the form of
where $\epsilon \in [0,1]$. Namely, the welfare-maximizing ITRs are determined by the signs of CATEs weighted by the share of corresponding preference type. For later comparison, we refer to the ITR satisfying (ref) as the naive ITR.
We have formalized the treatment assignment problem, presuming the true preference is observable. Practically, the true preference type is an unobservable feature. Instead, the policymaker must rely on the stated preference type by asking each individual about their preferred treatment. Based on the stated preference type, the policymaker determines the treatment for each individual according to the prespecified ITR. Generally speaking, the true and stated preferences do not necessarily concur. On the contrary, when individuals know the ITR before the preference statement, some have a strong incentive to make a false preference statement as exemplified in (ref).
We explicitly distinguish between the true and stated preferences to discuss the welfare-maximizing ITR in the presence of individuals' strategic revelation of preference type. Let $S_i(\delta) \in \mathcal{T}$ be individual $i$'s stated preference type when $i$ knows the applied ITR is $\delta$. The stated preference type is allowed to differ from the true preference type. Note that the stated preference type is a function of ITRs, implying that the stated preference can differ depending on the ITR implemented. Then, the treatment assigned is drawn from the Bernoulli distribution with parameter $\delta(S_i(\delta))$. We refer to this circumstance as environment under stated preference observation. As in the environment under true preference observation, the primal goal of the policymaker is to maximize the welfare (i.e., expected outcome). In the current environment, the welfare under an ITR $\delta$ is
Comparing (ref), observe that $W_{\operatorname{T}}$ is modified by replacing the true preference, $T_i$, with the stated preference, $S_i(\delta)$. The two welfare functions generally disagree as $S_i(\delta)$ does not necessarily correspond to $T_i$. We refer to the ITR maximizing $W_S$ as the optimal ITR. To proceed, we assume that the preference for treatments is naturally extended to the preference for lotteries over treatments ((ref)). This assumption allows us to describe when an individual makes the true preference statement. In its statement, a lottery over treatments is a vector $(p_1,p_0) \in [0,1]^2$ such that $p_1 + p_0 = 1$, where $p_d$ denotes the probability of getting treatment $d$.
This assumption says that each individual prefers the lottery that gives their preferred treatment with a higher probability. Hence, the preference for lotteries, characterized by (ref), is a reasonable extension of the preference for treatments. We can easily observe that the condition (ref) holds if and only if $\mathbb{E}_{d \sim \mathrm{Ber}(p_1)}[u_i(d)] \geq \mathbb{E}_{d \sim \mathrm{Ber}(q_1)}[u_i(d)]$ for any utility function $u_i: \mathcal{D} \to \mathbb{R}$ representing $\succsim_i$\footnote{A real-valued function $u_i:\mathcal{D} \to \mathbb{R}$ is said to be a utility function representing $\succsim_i$ if and only if $d \succsim_i d'$ is equivalent to $u_i(d) \geq u_i(d')$ for any pair $(d,d')$ of treatments.}. Thus, an alternative interpretation of (ref) is that each individual maximizes their expected utility. This assumption is common in studies of matching markets Kojima2010, Erdil2008, Erdil2014. With a slight abuse of notation, we write $(p_1,p_0) \succsim_i [\succ_i] ~ (q_1,q_0)$ when individual $i$ [strictly] prefers $(p_1,p_0)$ to $(q_1,q_0)$. Note that two lotteries are the same if and only if any individual is indifferent between the two lotteries. Given the preference for lotteries over treatments, we can discuss whether an ITR incentivizes the true preference revelation.
Under the strategy-proof ITR, each individual can obtain the lottery with (weakly) higher expected utility by telling the truth. In other words, any individual does not have a strong incentive to tell a lie in the preference statement. Moreover, the true preference revelation becomes the unique optimal behavior under the strictly strategy-proof ITRs. The following lemma gives a key to characterize strategy-proof ITRs.
(ref) is helpful for checking whether an ITR is strategy-proof. An ITR is strategy-proof precisely when it gives treatment 1 to individuals whose stated preference type is 1 with a higher probability than individuals whose stated preference type is 0. The reason is apparent: because $\delta(1) \geq \delta(0)$, individuals preferring treatment 1 can get their preferred treatment with a higher probability by telling the truth. The inequality is equivalent to $1-\delta(0) \geq 1-\delta(1)$; thus, the above interpretation also holds for individuals desiring treatment 0. (ref) allows us to characterize the stated preference as follows:
for each $i$. That is, the true and stated preferences agree [disagree] for any individual when $\delta(1) > [<] ~ \delta(0)$. Note that the two lotteries under the true and stated preferences statements are the same when $\delta(1) = \delta(0)$. Therefore, all individuals are indifferent between the true and false preference statements. In this case, the stated preference can be arbitrarily chosen. The behavior described in (ref) also yields a tractable representation of the welfare function in the environment under stated preference observation. Specifically, plugging (ref) into (ref) by cases, the welfare $W_{\operatorname{S}}(\delta)$ under an ITR $\delta$ equals
where equality follows from the law of iterated expectations. Note that the difference in the first and third cases in (ref) is that the role of $\delta(1)$ and $\delta(0)$ are swapped. Comparison of the expansions of the two welfare functions given in (ref) makes clear when the welfare functions in the environment under true and stated preference observation are different. The two welfare functions disagree when the false preference revelation is the unique optimal behavior.
As illustrated in (ref), the ITRs optimized ignoring individuals' strategic preference statements can lead to significant welfare losses. Then, the natural question is what kind of ITRs attain the highest welfare in the environment under stated preference observation. Moreover, are welfare maximization and strategy-proofness compatible? We answer these questions by deriving the oracle ITR under stated preference observation.
(ref) gives the optimal ITR $\delta^*$ under stated preference observation. This result yields three findings. First, the knowledge of $\beta_1 = \mathbb{P}(T_i=1)\tau(1)$, $\beta_0=\mathbb{P}(T_i=0)\tau(0)$, and $\beta_1 + \beta_0 = \mathbb{E}[Y_i(1) - Y_i(0)]$ are sufficient to construct the optimal ITR. In other words, it is sufficient to know the distribution of true preference type and the CATEs given the true preference type. The identification and estimation of the information will be discussed in (ref). Second, the naive and optimal ITRs are different. To understand how individuals' strategic preference statements induce the difference, we construct (ref). (ref) compares the naive and optimal ITRs given in (ref), by cases defined by the feature of the joint distribution of $(Y_i(0), Y_i(1), T_i)$. The first three columns show the signs of $\beta_1$, $\beta_0$, and $\beta_1 + \beta_0$. When the sign of $\beta_1 + \beta_0$ is implied by the signs of $\beta_1$ and $\beta_0$ or does not affect the structure of the oracle ITRs, the corresponding cell is left empty. To highlight the essential difference between the ITRs, the naive ITR is adjusted when its elements can be arbitrarily chosen from the unit interval to minimize the difference between the two ITRs. For instance, when $\beta_1 > 0$ and $\beta_0 = 0$, the naive ITR $\delta$ is given by $(\delta(1),\delta(0)) = (1,\eta)$ for arbitrary $\eta \in [0,1]$. In contrast, the optimal ITR is $(\delta^*(1),\delta^*(0)) = (1,1)$. In this case, we set $\eta = 1$ to make the two ITRs identical. When $\beta_1 = \beta_0 = 0$ or when $\beta_1 < 0 < \beta_0$ and $\beta_1 + \beta_0 = 0$, $\epsilon \in [0,1]$ can be arbitrarily chosen. Inspection of (ref) reveals that the essential difference between the two ITRs exists precisely when $\beta_1 < 0$ and $\beta_0 > 0$. In this case, a policymaker who ignores individuals' strategic preference revelation will try to assign treatment $1$ only to individuals who genuinely prefer treatment $0$. However, each individual can gain by lying about their preferred treatment. As a result, the individuals receiving treatment 1 are precisely the opposite of those the policymaker originally aimed at. Instead, (ref) implies that assigning the same treatment uniformly to all individuals regardless of the stated preference type maximizes welfare. The uniform treatment is determined by the sign of the average treatment effect, $\beta_1 + \beta_0 = \mathbb{E}[Y_i(1)-Y_i(0)]$. Last, the optimal ITR is always strategy-proof: no individual has a strong incentive for false preference revelation under the optimal ITR. This is obvious from (ref), since $\delta^*(1) \geq \delta^*(0)$ holds for any case. Moreover, $\delta^*(1)$ and $\delta^*(0)$ are equal except for the case when $\beta_1 > 0$ and $\beta_0 < 0$. In other words, individuals are indifferent between the two lotteries induced by the optimal ITR. Hence, individuals choose stated preferences arbitrarily. Nevertheless, this does not affect the welfare because the optimal ITR does not individualize the assignment. In contrast, the truthful preference revelation becomes the unique optimal behavior for all individuals when $\beta_1 > 0$ and $\beta_0 < 0$.
In (ref), we assumed that the policymaker knows the distribution of the true preference type, $\mathbb{P}(T_i=1)$, and the average treatment effect conditional on the true preference type, $\tau(t) = \mathbb{E}[Y_i(1) - Y_i(0)| T_i = t]$. Practically, these objects are unknown and should be identified and estimated from data. In this section, we introduce two particular experiment designs that allow us to identify $\mathbb{P}(T_i=1)$ and $\tau(t)$. Specifically, (ref) defines the strictly strategy-proof randomized controlled trial (SSP-RCT), an adjustment of the RCT with a pre-treatment survey, so that the true preference revelation becomes the strictly optimal behavior for any individual. (ref) discusses the doubly randomized preference trial (DRPT) Rucker1989,Wennberg1993. The DRPT randomly assigns individuals to treatment 0, treatment 1, and free-choice groups; in the former two groups, the treatment exposed is exogenously determined, while it is determined by individuals' choice in the third group. We demonstrate that both experimental designs can identify the objects of interest. Building on the identification of the key quantities, we develop data-dependent procedures to determine an ITR, presuming data derived from data generating processes like the SSP-RCTs or DRPTs. Concretely, we construct the statistical treatment rule (STR), a function that maps each possible realization of data to an ITR. Following Manski2004, we evaluate the performance of our proposed STRs based on the maximum regret. Formally, given a class $\mathcal{P}$ of data generating processes and an STR $\widehat{\delta}$, the maximum regret of the STR is given by
The expectation corresponds to the regret, the average loss from the use of $\widehat{\delta}$ relative to the highest welfare achievable when the true data generating process $P$ is known. Then, the maximum regret is defined by taking the supremum of the regret over the class of the data generating processes. The class $\mathcal{P}$ will be specified below. We derive the finite-sample upper bound of the maximum regret of our proposed STR. These results imply that the worst-case regret converges to zero at rate $n^{-1/2}$. In the following analysis, we suppose that the sample population is the same as the population $(I,\Sigma,\mathbb{P})$ of interest. \footnote{Generally, the sample population can differ from the population of interest as long as the joint distribution of $(Y_i(1),Y_i(0),T_i)$ is the same between the two populations and individuals of the experimental population maximizes their own expected utility.} Thus, each member $i$ of the sample population has potential outcomes, $Y_i(0)$ and $Y_i(1)$, and the true preference type, $T_i$, and (ref) is satisfied.
An idea of the strictly strategy-proof randomized controlled trial (SSP-RCT) is to adjust the propensity score of the RCT with a pre-treatment survey so that the true preference statement becomes the strictly optimal behavior for each individual. The trick to induce the true preference revelation comes from the observations in (ref). To be specific, consider an propensity score function $p:\mathcal{T}\to[0,1]$ such that
With this propensity score function being announced, each individual reports $S_i = S_i(p)$ in the pre-treatment survey. Then, each individual's experimental exposure $D_i \in \mathcal{D}$ is drawn from the Bernoulli distribution with parameter $p(S_i)$, and the outcome, $Y_i$, is observed according to $Y_i = Y_i(D_i)$. Thus, the observable data consists of $Y_i$, $D_i$, and $S_i$ for each $i$. Most importantly, condition (ref) ensures that the true preference statement becomes the utility-maximizing behavior (see (ref)). Therefore, we have $S_i = T_i$ for any individual $i$. In addition, the unconfoundedness holds by construction; that is, $(Y_i(0),Y_i(1)) \perp D_i \mid S_i$. As a result, this experimental design can identify $\mathbb{P}(T_i = 1)$ and $\tau(t)$. Specifically, it can be easily shown that
for any $t$ in the support of $T_i$. It is natural to ask whether observational studies containing stated preferences make the identification possible. The joint distributions of $(Y_i(0),Y_i(1),T_i,S_i,D_i)$ satisfying (ref) are sufficient for the identification, given that the observable data consists of $Y_i = Y_i(D_i)$, $D_i$ and $S_i$.
(ref) are standard in the study of statistical treatment rules Kitagawa2018,Mbakop2021,Zhou2023. (ref) can be weakened to the existence of expectations, $\mathbb{E}[Y_i(1)]$ and $\mathbb{E}[Y_i(0)]$, for the identification. We include this assumption only for the regret analysis below. (ref) requires that the true and stated preferences coincide for each individual. As illustrated above, this is satisfied if the joint distribution is induced by an SSP-RCT and (ref) holds. However, if we focus only on the satisfaction of (ref), this is possibly achieved by other methods. For instance, the literature on matching markets has developed strategy-proof assignment mechanisms Roth1982,Dubins1981,Ergin2002. For any distribution with (ref), the identification of $\mathbb{P}(T_i = 1)$ and $\tau(t)$ can be conducted in the same way as (ref). For fixed $M$ and $\kappa$, we denote by $\mathcal{P}_{\operatorname{SSP-RCT}}(M,\kappa)$ the class of joint distributions satisfying (ref) because the SSP-RCTs particularly meet this assumption. Now, we propose the STR that maps the data generated from the joint distribution in $\mathcal{P}_{\operatorname{SSP-RCT}}(M,\kappa)$ to an ITR. Suppose that we obtain $n$ iid draws from the joint distribution of $(Y_i(0),Y_i(1),T_i,S_i,D_i)$ and observe data $\{(Y_i,D_i,S_i)\}_{i = 1}^n$, where $Y_i = Y_i(D_i)$. Given this data, $\mathbb{P}(T_i = t) \tau(t)$ can be unbiasedly estimated by
We assume that the propensity score, $\mathbb{P}(D_i=1|S_i=t)$, is known. Our proposed STR, $\widehat{\delta}_{\operatorname{SSP-RCT}}$, is defined by replacing $\beta_t$ in (ref) with $\widehat{\beta}_t$. For simplicity, we set $(\widehat{\delta}_{\operatorname{SSP-RCT}}(1),\widehat{\delta}_{\operatorname{SSP-RCT}}(0)) = (0,0)$ when $\widehat{\beta}_1 = \widehat{\beta}_0 = 0$ or when $\widehat{\beta}_1 < 0 < \widehat{\beta}_0$ and $\widehat{\beta}_1 + \widehat{\beta}_0 = 0$. The following result gives the statistical performance of our STR in terms of the maximum regret.
(ref) provides the finite-sample upper bound of the maximum regret of the proposed STR $\widehat{\delta}_{\operatorname{SSP-RCT}}$. Whatever joint distribution of $(Y_i(0),Y_i(1),T_i)$ the population has, the maximum regret of the STR converges to zero at rate $n^{-1/2}$ as long as the data comes from the data generating process meeting (ref).
Doubly randomized preference trials (DRPTs) randomly assign individuals to three experimental groups: treatment 0, treatment 1, and free choice groups Rucker1989,Wennberg1993. The exposed treatment is exogenously determined in the former two groups, and non-compliance is not allowed. Specifically, treatment $d$ is given in the treatment $d$ group. In contrast, each individual in the free-choice group freely chooses their preferred treatment. At first glance, the DRPT may seem a sole extension of the classical RCT with two treatment groups. However, the existence of the free-choice group, combined with an additional assumption, allows us to identify $\mathbb{P}(T_i = 1)$ and $\tau(t)$. We first introduce some variables to describe the DRPT formally. For ease of exposition, we denote the treatment 0, treatment 1, and choice group by 0, 1, and 2, respectively, and let $\mathcal{Z} = \{0,1,2\}$ be the set of the experimental groups. On top of $Y_i(0)$ and $Y_i(1)$, suppose that individual $i$ has a potential outcome $Y_i(d,z)$ that would be realized if $i$ were assigned to group $z$ and exposed to treatment $d$. For each $z \in \mathcal{Z}$, let $D_i(z) \in \mathcal{D}$ be the potential treatment that individual $i$ would choose if $i$ was assigned to group $z$. As non-compliance is not allowed in the treatment 0 and 1 groups, we have $D_i(0) = 0$ and $D_i(1) = 1$ for all $i$. In the choice group, individuals choose the treatment according to their own preferences, whence $D_i(2) = T_i$ for each $i$. The DRPT determines the group to which $i$ belongs, $Z_i \in \mathcal{Z}$, by drawing a lottery over experimental groups. Then, the observable data consists of $Z_i$, the observed treatment $D_i = D_i(Z_i)$, and the observed outcome $Y_i = Y_i(D_i(Z_i),Z_i)$. By construction, the potential outcomes, potential treatments, and the true preference type are jointly independent of the assigned group; that is, $((Y_i(d,z))_{d\in\mathcal{D},z\in\mathcal{Z}},(D_i(z))_{z\in\mathcal{Z}}) \perp Z_i$. In DRPTs, the key assumption for the identification is the well-known exclusion restriction Angrist1996. That is,
This requires that whether the treatment exposed is determined exogenously or by their own choice does not affect the outcome. This assumption is often controversial in practice, but some methods exist to test its necessary condition. Specifically, Kitagawa2015 provides a statistical test for the necessary condition of assumptions required to identify the local average treatment effect, that is, the random assignment of $Z_i$, monotonicity, and exclusion restriction Angrist1996. In DRPTs, the former two assumptions are automatically satisfied by construction, and hence, the procedure tests the necessary condition of the exclusion restriction. Alternatively, the discussion in Section 7 of Long2008 suggests a test feasible under a particular experimental design that combines the strictly strategy-proof RCT and DRPT. Under the exclusion restriction, the DRPT can be used to identify $\mathbb{P}(T_i = 1)$ and $\tau(t)$ by viewing the assigned group $Z_i$ as the multi-valued instrumental variable Ida2022,Wing2017. First of all, $D_i(2) = T_i$ and random assignment of $Z_i$ implies
Because $D_i(2) = T_i \geq 0 = D_i(0)$, the CATE for individuals preferring treatment 1 is equivalent to the local average treatment effect (LATE) for individuals switching treatment as the instrument $z$ is exogenously changed from $0$ to $2$. More explicitly, we have
Given this connection, the results in Imbens1994,Angrist1996 imply that $\tau(1)$ can be identified as in
Similarly, the CATE for individuals preferring treatment 0 is the same as the LATE for individuals changing treatment as the exogenous switch of the instrument $z$ goes from $2$ to $1$. Hence, it follows that
With the interpretation of $Z_i$ as an instrument, (ref) is sufficient for the identification of $\mathbb{P}(T_i = 1)$ and $\tau(t)$ using the instrumental variable approach described above.
(ref) are parallel to (ref) in (ref). (ref) requires that the instrument creates groups under which the exposed treatments are determined exogenously and a group in which individuals freely choose according to their preference. The joint distribution induced by the DRPT fulfills this requirement. Under the joint distribution satisfying (ref), $\mathbb{P}(T_i = 1)$ and $\tau(t)$ are identified in the same manner as above. We denote the class of joint distributions with (ref) by $\mathcal{P}_{\operatorname{DRPT}}(M,\kappa)$ for fixed $M$ and $\kappa$ because DRPTs satisfy the assumption. We now propose the STR mapping the data generated from the data generating processes satisfying (ref) to an ITR. Suppose that we obtain $n$ iid draws from the joint distribution in $\mathcal{P}_{\operatorname{DRPT}}(M,\kappa)$ and observe data, $\{(Y_i,D_i,Z_i)\}_{i = 1}^n$ following $D_i = D_i(Z_i)$ and $Y_i = Y_i(D_i(Z_i),Z_i)$. Given this data, one can unbiasedly estimate $\beta_t = \mathbb{P}(T_i = t) \tau(t),t\in\mathcal{T}$ by
Again, the probabilities of group assignment, $\mathbb{P}(Z_i = z),z\in\mathcal{Z}$, are assumed to be known. This supposition is reasonable when the data is obtained from the DRPT. One can view $\widehat{\beta}_1$ and $\widehat{\beta}_0$ as unbiased estimators for the intention-to-treat effects. Then, our proposed STR, $\widehat{\delta}_{\operatorname{DRPT}}$, is defined by substituting $\widehat{\beta}_t$ for $\beta_t$ in (ref). For simplicity, we set $(\widehat{\delta}_{\operatorname{DRPT}}(1),\widehat{\delta}_{\operatorname{DRPT}}(0)) = (0,0)$ when $\widehat{\beta}_0 = \widehat{\beta}_1 = 0$ or when $\widehat{\beta}_1 < 0 < \widehat{\beta}_0$ and $\widehat{\beta}_1 + \widehat{\beta}_0 = 0$. The next result gives an upper bound of the finite-sample maximum regret of $\widehat{\delta}_{\operatorname{DRPT}}$.
(ref) ensures that the maximum regret of $\widehat{\delta}_{\operatorname{DRPT}}$ converges to zero at rate $n^{-1/2}$ as long as the data comes from the data generating process with (ref). This convergence rate is the same as that of $\widehat{\delta}_{\operatorname{SSP-RCT}}$ in (ref).
We demonstrate our proposed STR using the results reported in Wing2017. They analyzed data from a DRPT conducted with students in an introductory psychology class Clark2000. This DRPT examined the effect of vocabulary and mathematics training on test scores. The total number of participants in this DRPT was 450, and they were randomly assigned to one of the three groups: the vocabulary training group, mathematics training group, and free-choice group with probability $1/4$, $1/4$, and $1/2$, respectively. As a result, three experimental groups, the vocabulary training group, mathematics training group, and free-choice group, contained 116, 119, and 210 students, respectively. Fifty advanced vocabulary terms were taught in the vocabulary training group, while 5 algebraic concepts were taught in the mathematics training. In the following analysis, we regard vocabulary training as treatment 1 and mathematics training as treatment 0. Both treatments lasted about 15 minutes. After the training session, the participants took a post-test consisting of 30 vocabulary questions and 20 mathematics questions, regardless of which training was received. Of the 450 participants, 445 completed this experimental procedure. For a more detailed description of this experiment, see Shadish2008. (ref), adapted from Wing2017, shows the estimates of the share of the preferred treatment and the estimates of the CATEs on vocabulary and mathematics test scores given the preferred treatment. The estimates imply that $62\%$ of students prefer vocabulary training while $32\%$ prefer mathematics training. For students who preferred vocabulary learning, vocabulary learning improved vocabulary test scores by 8.5 points and reduced math test scores by 3.4 points compared to math learning. For students who preferred learning mathematics, vocabulary learning improved vocabulary test scores by 7.4 points and reduced mathematics test scores by 5.5 points compared to mathematics learning. All of the CATEs were significantly different from zero.
For illustrational purposes, we define the outcome of interest as the weighted sum of vocabulary and mathematics test scores. Formally, let $V_i(d)$ and $M_i(d)$ be the potential vocabulary and mathematics test scores under treatment $d$. Given a weight $w \in [0,1]$, the potential outcome of interest under treatment $d$ is defined by
The weight being equal to zero means we only care about the vocabulary test scores. As $w$ gets large, more emphasis is put on the mathematics test scores, and $w=1$ means that we focus only on the mathematics test scores. With this definition of the targeted outcome, we operate the STR proposed in (ref). (ref) draws determinants of ITR---$\widehat{\beta}_1$, $\widehat{\beta}_0$, and $\widehat{\beta}_1 + \widehat{\beta}_0$---by each weight of the targeted outcome. When the weight is less than $0.538$, both $\widehat{\beta}_1$ and $\widehat{\beta}_0$ are positive. Hence, our proposed STR indicates all students take the vocabulary training. When the weight is larger than $0.770$, both $\widehat{\beta}_1$ and $\widehat{\beta}_0$ is negative. In this case, the STR indicates all students take the mathematics training. When the weight is in $(0.538,0.770)$, $\widehat{\beta}_1$ is negative and $\widehat{\beta}_0$ is positive. At first glance, it would seem optimal to instruct those who prefer vocabulary training to learn math and those who prefer math training to learn vocabulary. However, upon learning of this ITR, students lie in their stated preferences. The resulting allocation achieved is not optimal. Instead, our STR does not personalize the assignment based on stated preferences but rather assigns the same training to all students. Specifically, when the weight is less than or equal to $0.637$, we assign vocabulary learning to all students; otherwise, we assign math learning to all students.
This study investigated the individualized treatment assignment problem based on stated preferences for treatments. When individuals know the deployed ITR before the preference statement, they strategically state their preferences. Under the assumption that individuals maximize their expected utility, we derived an optimal ITR that maximizes welfare. The optimal ITR is strategy-proof, that is, individuals have no strong incentive to make a false preference statement. The optimal ITR requires information about the distribution of the true treatment preference and the conditional average treatment effect given the true preference. We proposed two experimental designs---strictly strategy-proof RCTs (SSP-RCTs) and doubly randomized preference trials (DRPTs)---that allow us to identify the information. We developed statistical treatment rules, assuming that the data comes from either SSP-RCTs or DRPTs. The maximum regret of the proposed STRs converges to zero at a rate of the square root of the sample size. We focused on binary treatment assignment problems and have not mentioned the case of more than three treatments. It is easy to adapt the model developed in (ref) to accommodate more than three treatments. Here, we briefly demonstrate this for the case of three treatments. Each individual has a strict preference for the three treatments. Then, we can divide individuals into six types of preferences. Accordingly, the CATEs are defined for each preference type, and an ITR specifies a lottery over the treatments for each preference type. To extend the preference for treatments to the preference over lotteries, we can utilize the concept of first-order stochastic dominance Erdil2014,Erdil2008,Kojima2010. However, the form of the optimal ITR is unclear when three treatments exist. This is left for future research. Another issue not addressed in this study is an ethical one. When people have a preference for a treatment, is it ethical to give them a treatment that differs from their preferred treatment? In the medical context, this may be permissible as the physician often has more knowledge about the treatment than the patient and may be able to persuade the patient to accept the recommended treatment. However, this is not always permissible in public policy, and the pros and cons may vary depending on the context.