Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
109,538 characters · 16 sections · 32 citation commands
Shrinkage Methods for Treatment Choice
This study examines the problem of determining whether to treat individuals based on observed covariates. The most common decision rule is the conditional empirical success (CES) rule proposed by manski2004statistical, which is a rule assigning individuals to treatments that yield the best experimental outcomes conditional on the observed covariates. The CES rule uses only the average treatment effect (ATE) estimate conditional on each covariate value. By contrast, a common method in statistical estimation problems is to shrink unbiased but noisy preliminary estimates toward the average of these estimates. It is well known that shrinkage estimators may have smaller mean squared errors than unshrunk estimators. This study assumes that the dispersion of conditional ATEs (CATEs) is bounded and proposes a shrinkage rule that assigns individuals to treatments based on shrinkage estimators. We also propose a method to select the shrinkage factor by minimizing an upper bound of the maximum regret. By considering the treatment rules for individuals that are based not only on each CATE but also on the CATEs of others, it is possible to incorporate information across individuals. This allows the proposed shrinkage rule to perform as well as or better than existing treatment rules in the sense of maximum regret and to be more flexible to the heterogeneity of individuals. In addition, we compare the shrinkage rule with other rules when the space of CATEs is correctly specified or misspecified.
The contributions of this study are as follows. First, our approach is attractive from a computational perspective. The computation of the exact minimax regret rule is often challenging in the context of statistical treatment choice. Indeed, when the space of CATEs is restricted and the number of possible covariate values is large, it is difficult to obtain a shrinkage rule that minimizes the maximum regret. To overcome this problem, we propose a shrinkage rule that minimizes a tractable upper bound of the maximum regret. In this approach, each shrinkage factor is obtained by optimizing a single parameter and hence the proposed shrinkage rule is easy to compute.
Second, we compare the maximum regret of the shrinkage rule with those of alternative rules when the space of CATEs is correctly specified. As an alternative to the CES and shrinkage rules, one could consider using the pooling rule that determines whether to treat the individuals based on the average of the CATE estimates. As the CES and pooling rules are special cases of shrinkage rules, the proposed shrinkage rule is expected to outperform these two rules. However, because the proposed shrinkage rule does not minimize the exact maximum regret, its maximum regret may be larger than those of the CES and pooling rules. Therefore, it is essential to compare the maximum regrets. In Section (ref), we derive the conditions under which the proposed shrinkage rule outperforms the CES or pooling rules. If the dispersion of the CATEs is small compared with the standard deviations of the CATE estimates, then the maximum regret of the proposed shrinkage rule is less than that of the CES rule under homoscedasticity. A detailed definition of the dispersion of CATEs is provided in Section (ref). Intuitively, when the space of CATEs is small enough, the CES rule can be improved by shrinking each CATE estimate toward the average of these estimates. We also demonstrate that the shrinkage rule outperforms the pooling rule when the dispersion of the CATEs is sufficiently large. Furthermore, combined with these results, we show that the proposed shrinkage rule outperforms both the CES and pooling rules when the dispersion is moderate.
Third, we evaluate the maximum regret of the shrinkage rule when the space of CATEs is misspecified. The choice of the space of CATEs is important in practice because the minimax decision rule depends on the space of CATEs. For example, armstrong2018optimal,armstrong2021finite consider the minimax estimation and inference problem for treatment effects and show that it is not possible to choose the parameter space automatically in a data-driven manner. Hence, it is crucial to analyze the decision rule under the misspecification of the space of CATEs. We investigate the performance of the shrinkage rule and show that our results are robust to the misspecification of the space of CATEs. To the best of our knowledge, this is the first study to consider the misspecification of the parameter space in the treatment choice problem.
Consequently, this study contributes to the growing literature on statistical treatment choice initiated by manski2000identification, manski2004statistical. Following manski2004statistical, manski2007minimax, hirano2009asymptotics, stoye2009minimax, stoye2012minimax, and tetenov2012statistical, we focus on the maximum regrets of statistical treatment rules. Similar to stoye2012minimax, tetenov2012statistical, ishihara2021evidence, olea2023decision, and yata2021optimal, we assume that CATE estimates are normally distributed. This assumption can be approximately justified by the asymptotic normality.
The analysis in this study most closely relates to that of stoye2012minimax, who also considers Gaussian experiments for CATEs and investigates the properties of the minimax regret treatment rule when CATEs depend on covariates with bounded variations. stoye2012minimax shows that the CES rule achieves minimax regret when the dispersion of CATEs is sufficiently large, and the pooling rule achieves minimax regret when the dispersion is sufficiently small. These results imply that the CES rule is optimal when the CATEs vary significantly depending on the values of the covariates, and the pooling rule is optimal when the CATEs take almost the same values. However, we do not know the minimax regret rule when the dispersion is moderate. Given such circumstances, we propose a shrinkage treatment rule that includes both the CES and pooling rules as special cases and compare the shrinkage rule with the CES and pooling rules in terms of maximum regrets for any dispersion value.
The remainder of this paper is organized as follows. Section (ref) explains the decision problem and introduces the shrinkage rules. Section (ref) proposes a shrinkage rule that selects the shrinkage factor by minimizing the maximum regret's upper bound and analyzes the shrinkage rule's properties. Section (ref) compares the maximum regrets of shrinkage, CES, and pooling rules when the space of CATEs is correctly specified and misspecified. Section (ref) presents numerical analyses to compare the shrinkage rule with the CES and pooling rules. As an illustration, we apply our method to the experimental data from the National Job Training Partnership Act (JTPA) Study in Section (ref). Finally, Section (ref) concludes the paper.
Suppose that we have experimental data $\{(Y_i,D_i,X_i)\}_{i=1}^n$, where $X_i \in \{x_1, \ldots, x_K\}$ is a discrete covariate, $D_i \in \{0,1\}$ is a binary indicator of the treatment, and $Y_i$ is a post-treatment outcome. Suppose that $X_i$ represents a group and we want to determine whether to treat individuals in each group based on the data. For example, the group is determined by an individual's demographics, school, firm, or region. This setting is similar to that of manski2004statistical, who proposes the CES rule.
The CES rule assigns individuals to treatments that yield the best experimental outcomes conditional on covariates. For $k \in 1, \ldots, K$, we define
where $n_{d,k} \equiv \sum_{i=1}^n 1\{D_i = d, X_i = x_k\}$. Letting $Y_i(0), Y_i(1)$ be the potential outcomes, then $\theta_k$ can be interpreted as the ATE conditional on $x_k$, $E[Y_i(1)-Y_i(0)|X_i=x_k]$, under the unconfoundedness assumption $(Y_i(0),Y_i(1)) \mathop{\perp\!\!\!\perp} D_i | X_i$. In addition, $\hat{\theta}_k$ is a natural estimator of $\theta_k$. The CES rule determines whether to treat individuals with $x_k$ based on the sign of $\hat{\theta}_k$. Then, the treatment rules can be viewed as a map from estimates $\hat{\bm{\theta}} \equiv (\hat{\theta}_1, \ldots, \hat{\theta}_K)'$ to the binary decisions of the treatment choice. Hence, the CES rule can be expressed as follows:
Because $\hat{\theta}_k$ is consistent and asymptotically normal under some weak conditions, we assume that $\hat{\theta}_1, \ldots, \hat{\theta}_K$ are independently distributed and
where $\sigma_k$ is the standard deviation of $\hat{\theta}_k$. We assume that $\sigma_k$ is known. In practice, we can only construct a consistent estimator for $\sigma_k$. This assumption can be approximately justified by the asymptotic normality. If $\hat{\theta}_k$ has the asymptotic normality, the distribution of $\hat{\theta}_k$ is approximated by $N(\theta_k,\sigma_k^2)$. Because the treatment effect can vary with observable individual characteristics, we allow $\theta_k$ to vary across the covariates.
Given a treatment choice action $\bm{\delta} \equiv (\delta_1, \cdots, \delta_K)' \in \{0,1\}^K$, we define the welfare attained at $\bm{\delta}$ as follows:
where $\bm{\theta}=(\theta_1,\dots,\theta_K)'$, $p_k \equiv P(X = x_k)$ and $\mu_{d,k} \equiv E[Y| D=d, X = x_k]$ for $d \in \{0,1\}$ and $k=1, \ldots, K$. Note that $\theta_k$ is written as $\theta_k = \mu_{1,k} - \mu_{0,k}$. If we know the true value of $\bm{\theta}$, then the optimal treatment choice action is given by $$ \bm{\delta}^{\ast} \ \equiv \ \left( \delta_1^{\ast}, \cdots, \delta_K^{\ast} \right)' \ \equiv \ \left( 1\{\theta_1 \geq 0\}, \ldots, 1\{\theta_K \geq 0\} \right)'. $$ However, the treatment choice action $\bm{\delta}^{\ast}$ is infeasible because the true value of $\bm{\theta}$ is unknown.
Let $\hat{\bm{\delta}} : \mathbb{R}^K \to \{0,1\}^K$ be a treatment rule that maps the estimates $\hat{\bm{\theta}}$ to the binary decisions of treatment choice. The welfare regret of $\hat{\bm{\delta}}(\hat{\bm{\theta}}) \equiv \left( \hat{\delta}_1(\hat{\bm{\theta}}), \ldots, \hat{\delta}_K(\hat{\bm{\theta}}) \right)'$ is defined as
where $E_{\bm{\theta}}$ is the expectation with respect to the sampling distribution of estimates $\hat{\bm{\theta}}$ given the parameters $\bm{\theta}$. Following existing studies, we evaluate the treatment rule $\hat{\bm{\delta}}$ using the maximum regret \[ \max_{\bm{\theta} \in \Theta} R(\bm{\theta},\hat{\bm{\delta}}), \] where $\Theta$ is the space of $\bm{\theta}$. The minimax regret criterion selects the statistical treatment rule that minimizes the maximum regret.
The CES rule does not use $\hat{\theta}_l$ for $l \neq k$ to determine whether or not to treat individuals with $x_k$. However, in the problem of estimating $\bm{\theta} \equiv ( \theta_1, \ldots, \theta_K)'$, a common method is to shrink $\hat{\theta}_k$ toward the average of estimates $\mathrm{ave}(\hat{\bm{\theta}}) \equiv \frac{1}{K} \sum_{k=1}^K \hat{\theta}_k$ and it is well known that shrinkage estimators may have smaller mean squared errors than unshrunk estimators. Hence, we propose the following shrinkage rules $\hat{\bm{\delta}}^{\bm{w}} : \mathbb{R}^K \to \{0,1\}^K$ for $\bm{w} \equiv (w_1, \ldots, w_K)' \in [0,1]^K$.
where
When the vector of shrinkage factors $\bm{w}$ is $\bm{1} \equiv (1,\ldots,1)'$, the shrinkage rule $\hat{\bm{\delta}}^{\bm{w}}$ becomes the CES rule $\hat{\bm{\delta}}^{\text{CES}}$ defined in ((ref)). Hence, the class of shrinkage rules contains the CES rule as a special case. Furthermore, when $\bm{w}$ is $\bm{0} \equiv (0,\ldots,0)'$, this rule becomes the pooling rule $\hat{\bm{\delta}}^{\text{pool}}(\hat{\bm{\theta}}) \equiv \hat{\bm{\delta}}^{\bm{0}}(\hat{\bm{\theta}})$.
From ((ref)), we observe that
where $\overline{\theta} \equiv K^{-1}\sum_{k=1}^K \theta_k$ and $s_k^2(w_k)$ is the variance of $w_k \cdot \hat{\theta}_k + (1-w_k) \cdot \mathrm{ave}(\hat{\bm{\theta}})$, that is, $$ s_k^2(w_k) \ = \ \left\{ w_k^2 + 2w_k(1-w_k)/K \right\} \sigma_k^2 + (1-w_k)^2 \left\{ K^{-2} \sum_{k=1}^K \sigma_k^2 \right\}. $$ Hence, from ((ref)), the welfare regret of shrinkage treatment rule $\hat{\bm{\delta}}^{\bm{w}}(\hat{\bm{\theta}})$ can be written as follows:
where $\text{sgn}(x) \equiv 1\{x>0\} -1\{x<0\}$, $\Phi(\cdot)$ is the distribution function of $N(0,1)$, and the second equality follows from the symmetry of the normal distributions.
In this section, we consider how to choose the shrinkage factors $\bm{w}$ under the following assumption.
Under Assumption (ref), the space of CATEs $\bm{\theta}$ becomes
where the constant $\kappa$ can be interpreted as controlling the dispersion of parameters $\theta_{k}$ around the mean $\overline{\theta}$. This assumption is similar to Assumption 1 in stoye2012minimax. stoye2012minimax assumes that $| \mu_{d,k} - \mu_{d,l} | \leq \kappa$ for all $d \in \{0,1\}$ and $k,l \in \{1, \ldots, K\}$, where $\mu_{d,k} = E[Y|D=d,X=x_k]$. Because $\theta_k = \mu_{1,k} - \mu_{0,k}$, this assumption implies that
If Assumption (ref) holds, then we have $| \theta_k - \theta_l | \leq | \theta_k - \overline{\theta} | + | \theta_l - \overline{\theta} | \leq 2 \kappa$; thus ((ref)) is satisfied. Conversely, if ((ref)) holds, then Assumption (ref) is satisfied by replacing $\kappa$ with $2\kappa$. Because the shrinkage location is the average of $\hat{\bm{\theta}}$, we restrict the difference between $\theta_k$ and the average of $\bm{\theta}$. In Section (ref), we consider other shrinkage locations and impose the assumption that corresponds to the location.
The minimax regret criterion selects the shrinkage factors that minimize the maximum regret. From ((ref)), the optimal shrinkage factors are obtained by minimizing the following: \[ \max_{\bm{\theta} \in \Theta(\kappa)} \sum_{k=1}^K p_k \cdot \left\{ |\theta_k| \cdot \Phi \left( - \frac{|\theta_k| - (1-w_k) \cdot \text{sgn}(\theta_k) (\theta_k - \overline{\theta})}{s_k(w_k)} \right) \right\}. \] However, obtaining the optimal shrinkage factors becomes computationally challenging when $K$ is large. To overcome this problem, we propose selecting shrinkage factors that minimize an upper bound of the maximum regret. Because we have $\text{sgn}(\theta_k) (\theta_k - \overline{\theta}) \leq \kappa$ for any $\bm{\theta} \in \Theta(\kappa)$ and $\Phi(\cdot)$ is an increasing function, the regret of shrinkage rule $\hat{\bm{\delta}}^{\bm{w}}(\hat{\bm{\theta}})$ is bounded above by
Using this upper bound, we obtain
where $\eta(a) \equiv \max_{t \geq 0} \left\{ t \cdot \Phi(-t+a) \right\} = \max_{t \in \mathbb{R}} \left\{ |t| \cdot \Phi(-|t|+a) \right\}$ for $a \in \mathbb{R}$. tetenov2012statistical and ishihara2021evidence show that function $\eta(\cdot)$ is strictly increasing and convex. Figure (ref) displays the shape of $\eta(a)$. Using function $\eta(\cdot)$, we propose the following shrinkage factors:
where $\psi_k(w_k;\kappa) \equiv s_k(w_k) \eta \left( (1-w_k) \cdot \kappa/s_k(w_k) \right)$ for $k = 1, \ldots, K$. If the right-hand side of ((ref)) is a good approximation of the maximum regret, this rule is expected to be close to the minimax regret rule.
Our approach is attractive from a computational perspective. Indeed, the proposed shrinkage factors are easy to compute because ((ref)) is obtained by optimizing the objective function over a single parameter while it is difficult to obtain the shrinkage factors that minimize the exact maximum regret when the number of possible covariate values is large.
In this section, we illustrate the difference between the true maximum regret and the upper bound proposed in the previous section. For illustration, we consider the simple case where $\sigma_1 = \cdots = \sigma_K$ and $p_1 = \cdots = p_K$. In this case, the regret function $R(\bm{\theta},\hat{\bm{\delta}}^{\bm{w}})$ does not change when $(\theta_k,w_k)$ is replaced with $(\theta_j, w_j)$. This implies that a function $\bm{w} \mapsto \max_{\bm{\theta} \in \Theta(\kappa)} R(\bm{\theta},\hat{\bm{\delta}}^{\bm{w}})$ is permutation invariant because $\Theta(\kappa)$ is permutation invariant. In this setting, our upper bound is also permutation invariant and the proposed shrinkage rule satisfies $w_1^{\ast}(\kappa) = \cdots = w_K^{\ast}(\kappa)$. Hence, we focus on the shrinkage rules $\hat{\delta}^{\bm{w}}$ with $w_1 = \cdots = w_K$. Although we do not know that the minimax shrinkage rule satisfies $w_1 = \cdots = w_K$, it is natural to consider the shrinkage rules with $w_1 = \cdots = w_K$ because the maximum regret is permutation invariant. If we consider the shrinkage rule $\hat{\delta}^{(w,\ldots,w)}$, the upper bound ((ref)) can be written as \[ \overline{R}_{\mathrm{upper}}(w) \ \equiv \ \sum_{k=1}^K p_k \cdot \psi_k(w;\kappa) \ = \ \psi(w;\kappa), \ \ \ \text{for $w \in [0,1]$.} \] where $\psi(w;\kappa) \equiv s(w) \eta \left( (1-w)\kappa / s(w) \right)$ and $s(w) \equiv s_1(w) = \cdots = s_K(w)$. Similarly, we define $\overline{R}_{\mathrm{true}}(w) \equiv \max_{\bm{\theta} \in \Theta(\kappa)} R(\bm{\theta},\hat{\delta}^{(w,\ldots,w)})$ for $w \in [0,1]$. In the following, we compare the upper bound $\overline{R}_{\mathrm{upper}}(w)$ and the true maximum regret $\overline{R}_{\mathrm{true}}(w)$ both analytically and numerically.
In the following proposition, we show that the upper bound $\overline{R}_{\mathrm{upper}}(w)$ is less than almost twice the true maximum regret $\overline{R}_{\mathrm{true}}(w)$.
Proposition (ref) shows that $\overline{R}_{\mathrm{true}}(w) \leq \overline{R}_{\mathrm{upper}}(w) \leq 2 \overline{R}_{\mathrm{true}}(w)$ holds for all $w \in [0,1]$ when $K$ is even. In addition, we can obtain a tighter upper bound that depends on $w$ and $\kappa$. Although the upper bound $\frac{\overline{R}_{\mathrm{upper}}(w)}{L_{\mathrm{true}}(w)}$ is complex and may be difficult to interpret, we can show that it is equal to $1$ when $\kappa=0$. Hence, these results imply that our bound is relatively tight compared to the true maximum regret, especially when $\kappa$ is close to zero. As demonstrated later, the numerical evaluation also shows that the functional form of $\overline{R}_{\mathrm{upper}}(w)$ is similar to that of $\overline{R}_{\mathrm{true}}(w)$ and $\overline{R}_{\mathrm{upper}}(w)$ is exactly the same as $\overline{R}_{\mathrm{true}}(w)$ when $\kappa = 0$.
To obtain Proposition (ref), we need to derive a lower bound of $\overline{R}_{\mathrm{true}}(w)$. When $K$ is even, we have
where $$ \bm{\theta}_{t,\kappa} \ \equiv \ (\underbrace{t+\kappa, \ldots, t+\kappa}_{\text{$K/2$ elements}}, \underbrace{t-\kappa, \ldots, t-\kappa}_{\text{$K/2$ elements}})'. $$ Using this lower bound, we obtain the results of Proposition (ref). We expect that this parameterization provides a tighter lower bound because the difference between each element of $\bm{\theta}_{t,\kappa}$ and the average becomes $\pm \kappa$. In fact, in Section (ref), we obtain the upper bound ((ref)) by replacing $\theta_k - \overline{\theta}$ with $\pm \kappa$.
Figure (ref) shows the functional forms of $\overline{R}_{\mathrm{true}}(w)$ and $\overline{R}_{\mathrm{upper}}(w)$ for $K=20$ and $\kappa = 0, \, 0.25, \, 0.5, \, 0.75$.\footnote{We calculate $\overline{R}_{\mathrm{true}}(w)$ using the Monte Carlo approximation. We generate $\bm{\theta}_1, \ldots, \bm{\theta}_{n_{\mathrm{sim}}}$ from a distribution on $\Theta(\kappa)$ and approximate $\overline{R}_{\mathrm{true}}(w)$ as $\max_{i} R(\bm{\theta}_i,\hat{\delta}^{(w,\ldots,w)})$. Hence, $\overline{R}_{\mathrm{true}}(w)$ in Figure (ref) may be less than the true maximum regret function.} The solid and dashed lines denote $\overline{R}_{\mathrm{true}}(w)$ and $\overline{R}_{\mathrm{upper}}(w)$, respectively. In all settings, $\overline{R}_{\mathrm{true}}(w)$ and $\overline{R}_{\mathrm{upper}}(w)$ have similar functional forms and $\overline{R}_{\mathrm{true}}(1)=\overline{R}_{\mathrm{upper}}(1)$ holds. In particular, Figure (ref) shows that the upper bound $\overline{R}_{\mathrm{upper}}(w)$ is exactly equal to the true maximum regret $\overline{R}_{\mathrm{true}}(w)$ when $\kappa=0$.
Figure (ref) shows the ratio of $\overline{R}_{\mathrm{upper}}(w)$ to $\overline{R}_{\mathrm{true}}(w)$ for $K=20$ and $\kappa = 0, \, 0.25, \, 0.5, \, 0.75$. In all cases, the ratio $\overline{R}_{\mathrm{upper}}(w) / \overline{R}_{\mathrm{true}}(w)$ is less than 2, and these results are consistent with Proposition (ref). Specifically, even in the worst case, the ratio does not exceed 1.4 under our settings. We also find that the ratio tends to increase when the shrinkage factor $w$ falls within a certain range that varies depending on $\kappa$, however, Figure (ref) shows that the minimum value of $\overline{R}_{\mathrm{upper}}(w)$ is close to that of $\overline{R}_{\mathrm{true}}(w)$ for all $\kappa$. These results imply that the optimal rule based on the upper bound achieves near-optimal maximum regret. In Appendix B, we provide additional results on $\overline{R}_{\mathrm{true}}(w)$, $\overline{R}_{\mathrm{upper}}(w)$, and $\overline{R}_{\mathrm{upper}}(w) / \overline{R}_{\mathrm{true}}(w)$ for $K=4, \, 100$.
In this section, we focus on the balanced case where $\sigma_1 = \cdots = \sigma_K$ and $p_1 = \cdots = p_K$. Although this setting is unrealistic, it is an unfavorable setting for our shrinkage method. When constructing the upper bound ((ref)), we replace $\text{sgn}(\theta_k) (\theta_k - \overline{\theta})$ with its maximum value $\kappa$. Hence, our upper bound is equivalent to \[ \sum_{k=1}^K p_k \cdot \max_{\bm{\theta} \in \Theta(\kappa)} \left\{ |\theta_k| \cdot \Phi \left( - \frac{|\theta_k| - (1-w_k) \cdot \text{sgn}(\theta_k) (\theta_k - \overline{\theta})}{s_k(w_k)} \right) \right\}. \] Therefore, the difference between the true maximum regret and the upper bound is small when a small number of groups account for a large proportion of the population. Especially, when $p_k = 1$ for some $k$, the upper bound is equal to the true maximum regret.
In this section, we discuss the asymptotic behavior of the proposed shrinkage rule. We consider the following three asymptotic situations: (i) the dispersion of the CATEs becomes larger, that is, $\kappa \to \infty$, (ii) the dispersion of the CATEs becomes smaller, that is, $\kappa \to 0$, and (iii) the number of subgroups increases, that is, $K \rightarrow \infty$. As an example, suppose that a discrete covariate represents the region in which the experiment is conducted and we want to determine whether to treat individuals in each region. Since replacing $\sigma_1, \ldots, \sigma_K, \kappa$ with $c\sigma_1, \ldots, c\sigma_K, c\kappa$ does not change the value of $\psi_k(w_k;\kappa)$ for any $c > 0$, $\kappa \to \infty$ is equivalent to $\sigma_k \to 0$ for all $k$. Hence, situation (i) corresponds to the case where the sample size in each region goes to infinity but the dispersion of the region-specific treatment effects is fixed. Situation (ii) corresponds to the case where the dispersion of the region-specific treatment effects approaches zero. As discussed above, this situation is equivalent to the case where the standard errors $\sigma_1, \ldots, \sigma_K$ go to infinity. Situation (iii) corresponds to the case where the number of regions in which the experiment is conducted goes to infinity but the sample size in each region is fixed. In the shrinkage estimation literature, many studies focus on this type of asymptotic scenario.
First, we consider situation (i). Because $\eta(a)$ is convex, we have $\eta(a) \geq \eta(0) + \eta'(0) a$ for $a \geq 0$. This implies \[ \psi_k(w_k;\kappa) \ \geq \ \eta(0) \cdot s_k(w_k) + (1-w_k) \cdot \eta'(0) \kappa, \] where $\eta'(0) \simeq 0.226$ and the equality holds when $w_k = 1$. Hence, if the right-hand side is minimized at $w_k = 1$, we obtain $w_k^{\ast}(\kappa)=1$. Because the derivative of the right-hand side becomes \[ \eta(0) \cdot s_k'(w_k) - \eta'(0) \kappa, \] the shrinkage factor $w_k^{\ast}(\kappa)$ becomes one when $\eta(0) \cdot s_k'(w_k) \leq \eta'(0) \kappa$ holds for all $w_k \in [0,1]$. For $w_k \in [0,1]$, we obtain
As $s_k(w_k)$ is bounded away from zero, we obtain $w_k^{\ast}(\kappa)=1$ for a sufficiently large $\kappa$. Hence, the proposed shrinkage rule becomes the CES rule when $\kappa$ is sufficiently large.
Second, we consider situation (ii). As $\kappa \to 0$, we have that $$ \psi_k(w_k;\kappa) \ \to \ s_k(w_k) \eta(0), \ \ \text{for $w_k \in [0,1]$.} $$ Thus, the limit of $\psi_k(w_k; \kappa)$ is minimized at $w_k = \text{arg} \min_{w \in [0,1]} s_k^2(w)$. Hence, if the homoscedasticity assumption holds, that is, $\sigma_k = \sigma$ for all $k$, then the limit of $\psi_k(w_k;\kappa)$ is minimized at $w_k=0$. Hence, if the dispersion of the parameters decreases, the proposed shrinkage rule approaches the pooling rule.
If we consider $w_k \cdot \hat{\theta}_k + (1-w_k) \cdot \mathrm{ave}(\hat{\bm{\theta}})$ to be an estimator of $\theta_k$, then this estimator becomes unbiased when $w_k = 1$. Hence, these two asymptotic situations imply that the shrinkage factor $w_k^{\ast}(\kappa)$ chooses a less biased estimator when $\kappa$ is large and a small variance estimator when $\kappa$ is small. This result can be seen as a type of bias-variance trade-off.
Finally, we consider situation (iii). We assume $\frac{1}{K^2} \sum_{k=1}^K \sigma_k^2 \to 0$ as $K \to \infty$. Under this condition, $s_k(w_k)$ can be approximated as $w_k \sigma_k$. Hence, in this situation, we have $$ \psi_k(w_k;\kappa) \ \to \ \tilde{\psi}_k(w_k;\kappa) \ \equiv \ \sigma_k w_k \eta \left( (w_k^{-1} - 1) \cdot (\kappa / \sigma_k) \right). $$ By letting $\tilde{w}_k^{\ast}(\kappa) \equiv \text{arg} \min_{w_k \in [0,1]} \tilde{\psi}_k(w_k;\kappa)$, $w_k^{\ast}(\kappa)$ can be approximated by $\tilde{w}_k^{\ast}(\kappa)$ when $K$ is large. Hence, when $K$ is sufficiently large, $\tilde{w}_k^{\ast}(\kappa)$ is useful to understand the properties of $w_k^{\ast}(\kappa)$.
Proposition (ref) implies that we obtain results similar to above two situations even when $K$ is large. Specifically, the proposed shrinkage rule becomes the CES rule when $\kappa / \sigma_k$ is larger than approximately $3/4$ and $K$ is large, and the proposed shrinkage rule becomes the pooling rule when $\kappa = 0$ and $K$ is large.
In the previous sections, we consider the shrinkage rules that shrink toward the average of $\hat{\bm{\theta}}$. However, there are other options for shrinkage locations. In this section, we consider shrinkage rules that shrink toward other estimates such as a regression estimate, a weighted average estimate, and so on. Concretely, we consider the following shrinkage rules:
where $\tilde{\delta}^{w_k}_{k}(\hat{\bm{\theta}}) \equiv 1 \left\{ w_k \cdot \hat{\theta}_k + (1-w_k) \cdot \hat{\xi}_k \right\}$, $\hat{\xi}_k \equiv \sum_{j=1}^K \omega_{k,j} \hat{\theta}_j$, and $(\omega_{k,1}, \ldots, \omega_{k,K})'$ is a known weight vector. For example, if $\hat{\xi}_k$ is the regression estimate of $\hat{\theta}_k$ on a vector of regression variables $z_k$, then $\hat{\xi}_k$ can be written as \[ \hat{\xi}_k \ = \ z_k' \left( \bm{Z}'\bm{Z} \right)^{-1} \bm{Z}' \hat{\bm{\theta}} \ = \ \sum_{j=1}^K z_k' \left( \bm{Z}'\bm{Z} \right)^{-1} z_j \hat{\theta}_j, \] where $z_k$ can be different from $x_k$ and $\bm{Z} \equiv \left( z_1, \ldots, z_K \right)'$. If $X_i$ represents the region in which the experiment is conducted, then we can use region-specific characteristics as $z_k$. Similarly, we can also consider the shrinkage rule that shrinks toward the weighted average estimate $\hat{\xi}_k = \sum_{j=1}^K p_j \hat{\theta}_j$ or the weighted regression estimate $\hat{\xi}_k = z_k' \left( \sum_{j=1}^K p_j \cdot z_j z_j' \right)^{-1} \sum_{j=1}^K p_j \cdot z_j \hat{\theta}_j$.
Instead of Assumption (ref), we assume that the difference between $\theta_k$ and the estimand of $\hat{\xi}$ is bounded by $\kappa$.
If we consider shrinkage rules that shrink toward the regression estimate, then we obtain \[ \xi_k \ = \ z_k' \left( \bm{Z}'\bm{Z} \right)^{-1} \bm{Z}' \bm{\theta}, \] where $\xi_k$ is interpreted as a linear projection of $\theta_k$ on $z_k$. Hence, Assumption (ref) implies that residuals from regressing $\bm{\theta}$ on $\bm{Z}$ are bounded by $\kappa$. Assumption (ref) generalizes Assumption (ref) because we have $\xi_k = \sum_{j=1}^K \omega_{k,j} \theta_j = \overline{\bm{\theta}}$ when $\omega_{k,j} = 1/K$.
Similar to Section (ref), we propose selecting shrinkage factors that minimize an upper bound of the maximum regret under Assumption (ref). We observe that \[ w_k \cdot \hat{\theta}_k + (1-w_k) \cdot \hat{\xi}_k \ \sim \ N \left( w_k \cdot \theta_k + (1-w_k) \cdot \xi_k, \,\tilde{s}_k^2(w_k) \right), \] where $\tilde{s}_k^2(w_k)$ denotes the variance of $w_k \cdot \hat{\theta}_k + (1-w_k) \cdot \hat{\xi}_k$. Let $\Theta_{g}(\kappa)$ be the space of $\bm{\theta}$ satisfying Assumption (ref). Then, as in ((ref)), the maximum regret of the shrinkage rule ((ref)) is bounded by $$ \max_{\bm{\theta} \in \Theta_{g}(\kappa)} R(\bm{\theta}, \tilde{\bm{\delta}}^{\bm{w}}) \ \leq \ \sum_{k=1}^K p_k \cdot \tilde{s}_k(w_k) \eta \left( \frac{(1-w_k) \cdot \kappa}{\tilde{s}_k(w_k)} \right). $$ Hence, similar to $\bm{w}^{\ast}(\kappa)$, we propose the following shrinkage factors:
As discussed above, we can consider shrinkage rules that shrink toward a general location such as a regression estimate or a weighted average estimate. However, because analyzing the maximum regrets of such shrinkage rules is computationally extensive, we focus on the shrinkage rule proposed in Section (ref) in subsequent sections.
The proposed shrinkage rule does not minimize the maximum regret because $\bm{w}^{\ast}(\kappa)$ minimizes an upper bound of the maximum regret. Hence, it remains unclear whether the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}(\hat{\bm{\theta}})$ is smaller than that of $\hat{\bm{\delta}}^{\text{CES}}(\hat{\bm{\theta}})$ and $\hat{\bm{\delta}}^{\mathrm{pool}}(\hat{\bm{\theta}})$. This section compares the maximum regret of the proposed shrinkage rule with the CES and pooling rules when $\kappa$ is correctly specified or misspecified.
We compare the maximum regret of the proposed shrinkage rule with that of the CES and pooling rules when the true space of CATEs $\Theta(\kappa)$ is known. First, we compare the maximum regrets of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}(\hat{\bm{\theta}})$ and $\hat{\bm{\delta}}^{\text{CES}}(\hat{\bm{\theta}})$.
The first result of Theorem (ref) implies that the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ is less than or equal to that of $\hat{\bm{\delta}}^{\mathrm{CES}}$ when $\kappa \geq t^{\ast}(0) \cdot (\overline{\sigma} - \underline{\sigma}) \simeq 0.75 (\overline{\sigma} - \underline{\sigma})$ or $\kappa \leq \eta(0) \cdot ( \underline{\sigma} - s_0 ) \simeq 0.17 ( \underline{\sigma} - s_0 )$. This implies that the proposed shrinkage rule outperforms the CES rule under the homoscedasticity assumption $\sigma_1 = \cdots = \sigma_K = \sigma$. Even when the homoscedasticity assumption does not hold, the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ is less than or equal to that of $\hat{\bm{\delta}}^{\mathrm{CES}}$ for sufficiently small or large $\kappa$. For example, if $\overline{\sigma} = 1.5 \underline{\sigma}$, then we have \[ \max_{\bm{\theta} \in \Theta(\kappa)} R(\bm{\theta},\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}) \ \leq \ \max_{\bm{\theta} \in \Theta(\kappa)} R(\bm{\theta},\hat{\bm{\delta}}^{\mathrm{CES}}) \ \ \ \text{when $0.38 \underline{\sigma} \leq \kappa$ or $\kappa \leq 0.17 (\underline{\sigma} - s_0)$.} \]
The second result of Theorem (ref) implies that the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ is less than that of $\hat{\bm{\delta}}^{\mathrm{CES}}$. As discussed in Section (ref), the proposed shrinkage rule becomes the CES rule when $\kappa$ is sufficiently large. Because condition $\kappa < t^{\ast}(0) \cdot \left( 1-\frac{1}{K} \right) \overline{\sigma}$ implies $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)} \neq \hat{\bm{\delta}}^{\mathrm{CES}}$, the proposed shrinkage rule has a smaller maximum regret than the CES rule when $t^{\ast}(0) \cdot (\overline{\sigma} - \underline{\sigma}) \leq \kappa < t^{\ast}(0) \cdot \left( 1-\frac{1}{K} \right) \overline{\sigma}$. Hence, under the homoscedasticity assumption, the proposed shrinkage rule has a smaller maximum regret than the CES rule when $\kappa < t^{\ast}(0) \cdot \left( 1-\frac{1}{K} \right) \sigma \simeq 0.75 \left( 1-\frac{1}{K} \right) \sigma$. If the homoscedasticity assumption does not hold, it is unclear whether ((ref)) holds when $\kappa$ is between $0.75(\overline{\sigma} - \underline{\sigma})$ and $0.17(\underline{\sigma} - s_0)$. However, numerical simulations in Section (ref) show that ((ref)) holds for all $\kappa$ below a certain value in all correctly specified settings.
\if0
\fi
Next, we compare the maximum regrets of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}(\hat{\bm{\theta}})$ and $\hat{\bm{\delta}}^{\text{pool}}(\hat{\bm{\theta}})$.
The first result of Theorem (ref) implies that the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ is less than or equal to twice that of $\hat{\delta}^{\mathrm{pool}}$ for any $\kappa > 0$. In addition, the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ is less than or equal to that of $\hat{\delta}^{\mathrm{pool}}$ when $\kappa = 0$. Hence, the proposed shrinkage rule is not so inferior to the pooling rule when $\kappa$ is positive. If the CATEs are the same across the groups, that is, $\kappa = 0$, the proposed rule is superior to the pooling rule.
Because $\eta(\cdot)$ is strictly increasing, the second result of Theorem (ref) implies that the proposed shrinkage rule has a smaller maximum regret than the pooling rule when $\kappa$ is sufficiently large. If the homoscedasticity assumption holds, for any $\kappa > 0$, the condition $s_0 \eta(\kappa/s_0) > 2\eta(0) \left( \sum_{k=1}^K p_k \sigma_k \right)$ is equivalent to
where a function $\eta(a)/a$ satisfies $\eta(a)/a \geq 1/2$ and $\lim_{a \to \infty} \eta(a)/a = 1$. Hence, for any $K$, ((ref)) holds when $\kappa / \sigma$ is larger than $4 \eta(0) \simeq 0.68$. Furthermore, if $K$ goes to infinity, the condition ((ref)) becomes $\kappa / \sigma > 2 \eta(0) \simeq 0.34$.
Combining Theorems (ref) and (ref) yields conditions under which the proposed shrinkage rule is superior to both the CES and pooling rules.
Corollary (ref) implies that if $K$ is large, then the proposed shrinkage rule has a smaller maximum regret than both the CES and pooling rules when $\kappa$ is within a certain range. When $K$ is large, the first condition of ((ref)) is satisfied as long as the dispersion of standard errors is not too large. Because $\frac{\sum_{k=1}^K p_k \sigma_k}{\overline{\sigma}} \leq 1$ and $\frac{t^{\ast}(0) }{4 \eta(0)} \simeq 1.106$, the second condition of ((ref)) is satisfied when $K \geq 12$. Hence, both ((ref)) and ((ref)) hold for some $\kappa$ when $K$ is sufficiently large. Under the homoscedasticity assumption, the range ((ref)) becomes \[ 0.68 \sigma \ < \ \kappa \ < \ 0.75 \left( 1 - \frac{1}{K} \right) \sigma. \] The following remark shows that when the homoscedasticity assumption holds, the proposed shrinkage rule has a smaller maximum regret than both the CES and pooling rules over a wider range than ((ref)).
In the previous section, we assume that the space of CATEs $\Theta (\kappa)$ is known. However, in practice, it may be challenging to select a reasonable $\kappa$. In this section, we consider the case in which the researcher's choice of the space of CATEs $\Theta (\kappa')$ is different from the true space of CATEs $\Theta (\kappa)$, that is, $\kappa$ is misspecified.
First, we compare the maximum regret of the proposed shrinkage rule with that of the CES rule when $\kappa$ is misspecified.
In Theorem (ref), we use the parameter of CATEs $\Theta (\kappa')$, which may differ from the true space $\Theta(\kappa)$, to determine the shrinkage factors. The first result of Theorem (ref) implies that if $\kappa \leq \kappa'$, the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ is less than that of $\hat{\bm{\delta}}^{\mathrm{CES}}$ under conditions analogous to those in Theorem (ref). However, the range of $\kappa$ in which ((ref)) holds is narrower than the range of Theorem (ref) due to the misspecification of $\kappa$. The second result of Theorem (ref) implies that if $\kappa \geq \kappa'$, the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ is less than or equal to that of $\hat{\bm{\delta}}^{\mathrm{CES}}$ when $\kappa$ is sufficiently small. According to the proof of Theorem (ref), if $\kappa \leq \kappa'$, the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ is less than that of $\hat{\bm{\delta}}^{\mathrm{CES}}$ for all $\kappa$ under the homoscedasticity assumption. Conversely, if $\kappa > \kappa'$, the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ can be larger than that of $\hat{\bm{\delta}}^{\mathrm{CES}}$ even when the homoscedasticity assumption holds. In fact, numerical simulations in Section (ref) show that the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ is at most $1.7$ times less than that of $\hat{\bm{\delta}}^{\mathrm{CES}}$ for all $\kappa$ even when $\kappa$ is twice as large as $\kappa'$.
\if0 In Theorem (ref), we use the parameter space $\Theta (\kappa')$, which may differ from the true parameter space $\Theta(\kappa)$, to determine the shrinkage factors. Hence, the upper bound of ((ref)) differs from ((ref)). However, when $\kappa' = \kappa$, Theorems (ref) and (ref) are equivalent.
Theorem (ref) implies that we obtain an upper bound similar to that of Theorem (ref) even when $\kappa$ is misspecified. When $\overline{\sigma} - \underline{\sigma} \leq \kappa / t^{\ast}(0)$ holds, the upper bound becomes one when $\kappa' \geq \kappa$. This implies that the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ is not grater than that of $\hat{\bm{\delta}}^{\mathrm{CES}}$ when $\kappa' \geq \kappa$. Because $\eta(a) \geq a/2$ from Lemma (ref) in Appendix 1, we obtain $H(a) \equiv a/\eta(a) \leq 2$. This implies that when $\kappa' < \kappa$ and $\overline{\sigma} - \underline{\sigma} \leq \kappa / t^{\ast}(0)$, we have \[ \frac{\max_{\bm{\theta} \in \Theta(\kappa)} R(\bm{\theta},\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')})}{\max_{\bm{\theta} \in \Theta(\kappa)} R(\bm{\theta},\hat{\bm{\delta}}^{\mathrm{CES}})} \ \leq \ 1 + 2 \left( \frac{ \kappa - \kappa'}{\kappa'} \right). \] Hence, the upper bound is close to one if $(\kappa - \kappa')/\kappa'$ is close to zero. As discussed in Section 3.2, the shrinkage factor $w_k^{\ast}(\kappa')$ becomes one for a sufficiently large $\kappa'$. Hence, if $\overline{\sigma} - \underline{\sigma} \leq \kappa / t^{\ast}(0)$ holds and $\kappa'$ is sufficiently large, the upper bound becomes \[ 1 + H(0) \cdot \left( \frac{|\kappa - \kappa'|_{+}}{\kappa'} \right) \ = \ 1. \] This is because the proposed shrinkage rule $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ becomes the CES rule when $\kappa'$ is sufficiently large. \fi
Next, we compare the maximum regrets of the proposed shrinkage rule with the pooling rule when $\kappa$ is misspecified.
The first result of Theorem (ref) implies that if $\kappa \leq \kappa'$, the maximum regret of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ is less than that of $\hat{\bm{\delta}}^{\mathrm{pool}}$ under the same condition as Theorem (ref). The second result of Theorem (ref) also implies that ((ref)) holds when $\kappa$ is sufficiently large, but the range of $\kappa$ in which ((ref)) holds is narrower than the range of Theorem (ref) due to the misspecification of $\kappa$. Similar to Theorem (ref), the third result of Theorem (ref) implies that the proposed shrinkage rule is not so inferior to the pooling rule if the degree of misspecification is smaller.
If the homoscedasticity assumption holds and $\kappa' = (1+c) \kappa$, Theorems (ref) and (ref) imply that both ((ref)) and ((ref)) hold when $\kappa / \sigma$ satisfies \[ \frac{\eta^{-1} \left( 2 \eta(0) \sqrt{K} \right)}{\sqrt{K}} \ < \ \frac{\kappa}{\sigma} \ < \ \frac{t^{\ast}(0) \left( 1 - \frac{1}{K} \right)}{1+c}. \] Hence, although the above range is narrower than ((ref)), the proposed shrinkage rule has a smaller maximum regret than the CES and pooling rules for some $\kappa$ when $\kappa'$ is larger than $\kappa$. On the other hand, when $\kappa'$ is less than $\kappa$, Theorems (ref) and (ref) do not establish whether there exists $\kappa / \sigma$ such that both ((ref)) and ((ref)) hold. However, numerical simulations in Section (ref) show that both ((ref)) and ((ref)) hold for some $\kappa$ even when $\kappa'$ is half of $\kappa$.
This section provides sufficient conditions for the proposed shrinkage rule to have a smaller maximum regret than the CES and pooling rules. While these conditions are intuitive and easy to interpret, it is possible to improve these conditions. However, attempting to improve the result leads to sufficient conditions that are difficult to interpret. Therefore, in this section, we present less sharp but more interpretable conditions.
\if0 As in Theorem (ref), when $\kappa' = \kappa$, the upper bound of Theorem (ref) is the same as that of Theorem (ref). When $\kappa' \geq \kappa$, the bound of ((ref)) implies that:
Because $\eta'(a) \leq 1$ from Lemma (ref) in Appendix 1, we have $\eta(a) \leq \eta(a') + (a-a')$ for $a \geq a'$. Hence, when $\kappa' \geq \kappa$, the first bound is bounded by
The upper bound of ((ref)) approaches two if $(\kappa' - \kappa)/\kappa$ is close to zero even when $\kappa$ is misspecified. Similarly, the second bound is bounded by the following:
Hence, the upper bound of ((ref)) approaches one as $\kappa'$ and $\kappa$ approach zero. The third bound of ((ref)) is identical to that in Theorem (ref), implying that the maximum regret of the shrinkage rule is smaller than that of the pooling rule when $\kappa \geq 0.68 \left( \sum_{k=1}^K p_k \sigma_k \right)$ and the maximum regret of the pooling rule can be much larger than that of the proposed shrinkage rule.
Because $\psi_k(w_k;\kappa)$ is increasing in $\kappa$, when $\kappa' \leq \kappa$, we obtain
Hence, in this case, the ratio of the maximum regrets is bounded by bound ((ref)) up to the following term: \[ 1 + \max_{k} \left\{ H \left( \frac{(1-w_k^{\ast}(\kappa')) \cdot \kappa' }{s_k \left( w_k^{\ast}(\kappa') \right)} \right) \right\} \cdot \left( \frac{\kappa - \kappa'}{\kappa'} \right). \] This implies that the upper bound of Theorem (ref) is approximately equal to that of Theorem (ref) when $(\kappa - \kappa')/\kappa'$ approaches zero. \fi
We present numerical examples to illustrate the results obtained in the previous sections. We demonstrate the relationship between $\kappa$ and $w_k^{\ast}(\kappa)$ under the homoscedasticity assumption. When homoscedasticity holds, that is, $\sigma_1 = \cdots = \sigma_K = 1$, we obtain $w_1^{\ast}(\kappa) = \cdots = w_K^{\ast}(\kappa) = w^{\ast}(\kappa)$. We calculate the shrinkage factor $w^{\ast}(\kappa)$ numerically for each $\kappa \in [0,1]$. Figure (ref) shows the relationship between $\kappa$ and $w^{\ast}(\kappa)$ for $K = 2, \, 5, \, 100$. The proposed shrinkage rule becomes the CES rule when $\kappa$ is sufficiently large and approaches the pooling rule when $\kappa$ approaches zero in all settings. This finding is consistent with the results presented in Section (ref). Furthermore, as seen in Proposition (ref), when the number of subgroups increases ($K=100$), the shrinkage factor becomes one if $\kappa$ is larger than $t^{\ast}(0) \simeq 0.752$.
Figure (ref) also shows the shrinkage factor of the James-Stein-type estimator when $K=100$ and $\hat{\theta}_k \sim N((-1)^k \kappa, 1)$. Concretely, the dot-dashed line denotes the median of $\hat{w}_{\mathrm{JS}}$ in Remark (ref). In this setting, $\bm{\theta}$ is contained in $\Theta(\kappa)$ because $\theta_k = (-1)^k \kappa$ and $\overline{\theta} = 0$. Similar to $w^{\ast}(\kappa)$, the median of $\hat{w}_{\mathrm{JS}}$ increases as $\kappa$ increases. However, $\hat{w}_{\mathrm{JS}}$ tends to be smaller than $w_k^{\ast}(\kappa)$ for almost all $\kappa$. This implies that the proposed shrinkage rule places greater emphasis on the bias than on the variance compared with the James–Stein-type estimator.
We compare the maximum regrets of the shrinkage, CES, and pooling rules when $\kappa$ is correctly specified. We consider the following two cases:
Case (i) corresponds to the balanced design, where the standard errors and group sizes are the same for all groups. Case (ii) corresponds to the unbalanced design, where the group sizes of one half of the groups are larger than those of the other half and the standard errors for the larger groups are smaller than those for the remaining groups. Figure (ref) shows the maximum regrets of the shrinkage, CES, and pooling rules for $\kappa = 0, \, 0.1, \, \ldots, \, 0.9, \, 1.0$ when $K=10$ and $100$.\footnote{Similar to Figure (ref), we calculate the maximum regrets using the Monte Carlo approximation.} In Case (i), we also calculate the treatment rule based on the James-Stein-type estimator $\hat{\bm{\delta}}^{\mathrm{JS}}(\hat{\bm{\theta}}) \equiv \left( \hat{\delta}^{\mathrm{JS}}_1(\hat{\bm{\theta}}), \ldots , \hat{\delta}^{\mathrm{JS}}_K(\hat{\bm{\theta}}) \right)'$ defined as \[ \hat{\delta}^{\mathrm{JS}}_k(\hat{\bm{\theta}}) \ \equiv \ 1 \left\{ \hat{w}_{\mathrm{JS}} \cdot \hat{\theta}_k + (1-\hat{w}_{\mathrm{JS}}) \cdot \mathrm{ave}(\hat{\bm{\theta}}) \geq 0 \right\}. \] Because this James-Stein-type estimator is not theoretically justified under heteroscedasticity, we do not report $\hat{\bm{\delta}}^{\mathrm{JS}}$ in Case (ii).
Figures (ref) and (ref) show that the maximum regret of the proposed shrinkage rule is less than or equal to that of the CES rule for all $\kappa$ in the balanced design. This result is consistent with Theorem (ref). In the balanced design, the proposed shrinkage rule is slightly inferior to the pooling rule when $K=10$ and $\kappa = 0.1$, but the maximum regrets of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ are less than or equal to those of $\hat{\bm{\delta}}^{\mathrm{pool}}$ in all other settings. In addition, the maximum regrets of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ are less than that of $\hat{\bm{\delta}}^{\mathrm{JS}}$ in almost all settings.\footnote{When calculating the maximum regret of $\hat{\bm{\delta}}^{\mathrm{JS}}$, we also calculate the regret $R(\bm{\theta},\hat{\bm{\delta}}^{\mathrm{JS}})$ using the Monte Carlo approximation. In Figure (ref), the maximum regret of $\hat{\bm{\delta}}^{\mathrm{JS}}$ is less than that of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ for $\kappa = 1$. However, this is probably due to approximation errors.} However, the difference between $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ and $\hat{\bm{\delta}}^{\mathrm{JS}}$ decreases as $K$ increases. In the unbalanced design, similar results are obtained for the shrinkage, CES, and pooling rules. In Appendix B, we provide additional results for $K=2$ and $4$.
Next, we calculate the maximum regret of the shrinkage rule when $\kappa$ is misspecified, that is, we calculate $\max_{\bm{\theta} \in \Theta(\kappa)} R(\bm{\theta},\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')})$. We consider the four misspecification situations, $\kappa' = 0.5 \kappa$, $\kappa' = 0.8 \kappa$, $\kappa' = 1.2 \kappa$, and $\kappa' = 2 \kappa$. Figure (ref) shows the maximum regrets of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$, $\hat{\bm{\delta}}^{\mathrm{CES}}$, $\hat{\bm{\delta}}^{\mathrm{pool}}$, and $\hat{\bm{\delta}}^{\mathrm{JS}}$ in the balanced design when $K=10$ and $\kappa$ is misspecified. Because $\hat{\bm{\delta}}^{\mathrm{CES}}$, $\hat{\bm{\delta}}^{\mathrm{pool}}$, and $\hat{\bm{\delta}}^{\mathrm{JS}}$ is not affected by misspecification, the maximum regrets of $\hat{\bm{\delta}}^{\mathrm{CES}}$, $\hat{\bm{\delta}}^{\mathrm{pool}}$, and $\hat{\bm{\delta}}^{\mathrm{JS}}$ are the same as those in Figure (ref).
When $\kappa \leq \kappa'$, as discussed in Section (ref), the maximum regrets of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ are less than or equal to those of $\hat{\bm{\delta}}^{\mathrm{CES}}$ in all settings. However, the range of $\kappa$ in which $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ has a smaller maximum regret than $\hat{\bm{\delta}}^{\mathrm{CES}}$ is narrower than that of Figure(ref) due to the misspecification of $\kappa$. In addition, the proposed shrinkage rule has a smaller maximum regret than the CES and pooling rules for some $\kappa$ when the degree of misspecification is not so large.
When $\kappa \geq \kappa'$, as discussed in Section (ref), the maximum regrets of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ are less than those of $\hat{\bm{\delta}}^{\mathrm{pool}}$ when $\kappa$ is sufficiently large. In contrast to the case where $\kappa \leq \kappa'$, the maximum regrets of $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ are larger than those of $\hat{\bm{\delta}}^{\mathrm{CES}}$ for some $\kappa$. However, the amount by which $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ exceeds $\hat{\bm{\delta}}^{\mathrm{CES}}$ decreases as the degree of misspecification decreases. In summary, we obtain results similar to the correctly specified case when the degree of misspecification is small, whereas $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')}$ may be inferior to other rules when $\kappa'$ is significantly smaller than $\kappa$.
\if0 We consider simple settings where $K=2$, $(\sigma_1,\sigma_2) = (1,1), (0.75,1.25)$, and $(p_1,p_2)=(0.5,0.5), (0.75, 0.25)$ to compare the maximum regrets of the shrinkage, CES, and pooling rules. Figures (ref)-(ref) show the maximum regrets of the shrinkage, CES, and pooling rules when $\kappa$ is correctly specified. If $p_1 > p_2$, then the number of units with the covariate $x_1$ is expected to be larger than the number of units with $x_2$. Hence, the standard deviation of $\hat{\theta}_1$ is expected to be smaller than that of $\hat{\theta}_2$.
As expected from Theorem (ref), the maximum regret of the shrinkage rule is always less than or equal to that of the CES rule in all settings. Additionally, the maximum regret of the shrinkage rule is equal to that of the CES rule when $\kappa$ is large. This is because the shrinkage rule becomes a CES rule when $\kappa$ is sufficiently large. Although the pooling rule is better than the shrinkage rule for some $\kappa$, as expected from Theorem (ref), the shrinkage rule is not worse than the pooling rule when $\kappa$ is small. Additionally, as $\kappa$ increases, the maximum regret of the pooling rule increases.
Next, we calculate the maximum regret of the shrinkage rule when $\kappa$ is misspecified, that is, we calculate $\max_{\bm{\theta} \in \Theta(\kappa)} R(\bm{\theta},\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa')})$. We consider two cases: (1) $\kappa' = 1.2 \kappa$ and (2) $\kappa' = 0.8 \kappa$. In case (1), the researcher's choice of the space of CATEs is larger than the true space of CATEs. In case (2), the researcher's choice of the space of CATEs is smaller. Figures (ref)-(ref) show the maximum regrets of the shrinkage, CES, and pooling rules in case (1) and Figures (ref)-(ref) show the maximum regrets in case (2).
Even when $\kappa$ is misspecified, we obtain results similar to those shown in Figures (ref)--(ref). These results imply that our shrinkage rule is robust to the misspecification of $\kappa$. In all settings, the maximum regret of the shrinkage rule is always less than or equal to that of the CES rule. Theorem (ref) implies that the shrinkage rule is superior to the CES rule when $\kappa' \geq \kappa$. However, Figures (ref)-(ref) show that similar results are obtained in these settings, even when $\kappa' \leq \kappa$. In these numerical examples, the maximum regret of the shrinkage rule decreases when $\kappa'$ is smaller than $\kappa$. This implies that the proposed shrinkage factors might be too large in some settings, as the choice of shrinkage factor minimizes the upper bound of the maximum regret. \fi
We illustrate the proposed method by applying it to experimental data from the National Job Training Partnership Act (JTPA) Study and using the dataset in abadie2002instrumental. The JTPA study is a randomized controlled trial whose purpose was to measure the impact of a training program on earnings. It also collected background information on the applicants prior to the random assignment and obtained data on their earnings for 30 months following the assignment.
We construct 24 subgroups using the following characteristics: race (Black, Hispanic, or other), sex (male or female), marital status (married or unmarried), and working status prior to random assignment (worked for at least 12 weeks in the 12 months preceding random assignment or not). Table (ref) shows the corresponding group numbers. For each subgroup, we calculate the treatment effect of the training program on earnings for 30 months following the assignment and its standard error. Following kitagawa2018should, we set the treatment cost as \$774. Hence, we use the CATE estimate minus \$774 as $\hat{\theta}_k$. Figure (ref) shows that the benefits of the training program vary across subgroups but are statistically insignificant for all subgroups.
As the average of $\hat{\bm{\theta}}$ is approximately \$541, the pooling rule determines to treat the individuals in all subgroups. However, Figure (ref) indicates that the decisions of the CES rule vary across the subgroups. Hence, the shrinkage rule may differ from the CES rule depending on the values of the shrinkage factors.
In this empirical application, we consider two James-Stein-type rules. The first is a treatment rule based on the following James–Stein-type shrinkage estimator:
where \[ \hat{w}_k \ \equiv \ \max \left\{ 1 - \frac{(K-3) \sigma_k^2}{\sum_{j=1}^K \left( \hat{\theta}_k - \mathrm{ave}(\hat{\bm{\theta}}) \right)^2} , 0\right\}. \] In contrast to the James–Stein-type shrinkage estimator introduced in Remark (ref), this estimator is specifically designed to ensure that the shrinkage factor remains non-negative. Using ((ref)), we define the following James-Stein-type rule:
Because the shrinkage estimator ((ref)) is not theoretically justified under heteroscedasticity, we consider an alternative James-Stein-type rule that accommodates heteroscedasticity. xie2012sure propose the following James–Stein-type shrinkage estimator in the heteroscedastic normal means model:
where $\tilde{w}_{k} \ \equiv \ \frac{\hat{\lambda}}{\sigma_k^2 + \hat{\lambda}}$ and \[ \hat{\lambda} \ \equiv \ \mathrm{arg} \min_{\lambda \geq 0} \left[ \frac{1}{K} \sum_{k=1}^K \left( \frac{\sigma_k^2}{\sigma_k^2 + \lambda} \right)^2 \left\{\hat{\theta}_k - \mathrm{ave}(\hat{\bm{\theta}})\right\}^2 + \frac{1}{K} \sum_{k=1}^K\left( \frac{\sigma_k^2}{\sigma_k^2 + \lambda} \right) \left( \lambda - \sigma_k^2 + \frac{2}{K} \sigma_k^2 \right) \right]. \] xie2012sure establish the asymptotic optimality property for this shrinkage estimator when $K \to \infty$. Using this shrinkage estimator ((ref)), we define the following James-Stein-type rule:
We calculate the shrinkage factors $\bm{w}^{\ast}(\kappa)$ for $\kappa = 500, \, 1{,}000$. Figure (ref) shows the relationship between the shrinkage factor and standard error. As expected, the shrinkage factor decreases as the standard error increases. For $\kappa = 1{,}000$, the shrinkage rule becomes the CES rule when $\sigma_k$ is less than around $1{,}300$. Proposition (ref) indicates that $w^{\ast}_k (\kappa)$ approaches $1$ asymptotically when $\kappa / \sigma_k \leq t^{\ast}(0) \simeq 0.75$, which is equivalent to $\sigma_k < \kappa / t^{\ast}(0) \simeq 1.33 \kappa$. Hence, this result indicates that the approximation in Proposition (ref) is useful.
Figures (ref) and (ref) illustrate the shrinkage estimates $w_k^{\ast}(\kappa) \cdot \hat{\theta}_k + (1-w_k^{\ast}(\kappa)) \cdot \text{ave}(\hat{\bm{\theta}})$ for $\kappa = 500, \, 1{,}000$. As the red line denotes the average of $\hat{\bm{\theta}}$, the shrinkage estimate (white circle) is closer to the red line than $\hat{\theta}_k$ (black circle). As the average of $\hat{\bm{\theta}}$ is positive, the decision of the shrinkage rule is the same as that of the CES rule when $\hat{\theta}_k$ is positive. Whereas, the decision regarding the shrinkage rule can differ from that regarding the CES rule when $\hat{\theta}_k$ is negative. Figure (ref) shows that the decision of the shrinkage rule is not identical to that of the CES rule in certain subgroups. However, the shrinkage rule makes the same decisions as the CES rule when $\kappa = 1{,}000$.
In addition, we also compute the James-Stein-type rules $\hat{\bm{\delta}}^{\mathrm{JS}}$ and $\hat{\bm{\delta}}^{\mathrm{XKB}}$. Figures (ref) and (ref) illustrate the James-Stein-type estimates $\hat{\theta}^{\mathrm{JS}}_k$. These estimates are close to our shrinkage estimates $w_k^{\ast}(\kappa) \cdot \hat{\theta}_k + (1-w_k^{\ast}(\kappa)) \cdot \text{ave}(\hat{\bm{\theta}})$ when $\kappa = 500$. Moreover, $\hat{\bm{\delta}}^{\mathrm{JS}}$ makes the same decisions as the shrinkage rule $\hat{\bm{\delta}}^{\bm{w}^{\ast}(\kappa)}$ for $\kappa = 500$. In contrast, all shrinkage factors of ((ref)) are approximately zero. This implies that $\hat{\bm{\delta}}^{\mathrm{XKB}}$ makes the same decision as the pooling rule. The pooling rule is nearly optimal when $\kappa$ is small and, as we will show later, the hypothesis $H_0:\bm{\theta} \in \Theta(0)$ is not rejected. Hence, although this result may appear extreme, it is not necessarily unreasonable. We note that $\hat{\theta}^{\mathrm{XKB}}_k$ is asymptotically justified but may be unreliable when $K$ is small.
This analysis focuses on the treatment choice problem when $\kappa = 500, \, 1{,}000$. As $w^{\ast}_k(\kappa)$ is increasing with respect to $\kappa$, the shrinkage rule makes the same decisions as the CES rule when $\kappa$ is greater than $1{,}000$. Hence, the decision regarding the shrinkage rule differs from that regarding the CES rule only when $\kappa$ is small. However, the choice of $\kappa = 500$ is not unrealistic. As discussed in Remark (ref), we can assess whether a given value of $\kappa$ is too small. Letting $Z_k \sim N(0,\sigma_k^2)$, the $95 \%$ quantile of $\max_{1\leq k \leq K} |Z_k - \overline{Z}|$ is about $9{,}500$. Whereas, the realized value of $\max_{1 \leq k \leq K} \left| \hat{\theta}_k - \mathrm{ave}(\hat{\bm{\theta}}) \right|$ is $4{,}256$. This implies that the hypothesis $H_0:\bm{\theta} \in \Theta(0)$ is not rejected and any value of $\kappa$ is consistent with the actual data. Therefore, this empirical application shows that the decision of the shrinkage rule can differ from that of the CES rule even if $\kappa$ is a realistic value.
Finally, we compare the proposed shrinkage rule with the empirical welfare maximization method. As discussed in Remark (ref), the EWM rule is equivalent to the CES rule if there are no restrictions on the class of candidate treatment rules. In this analysis, we consider the constraints that decisions are not changed based on race and that if men with the same characteristics receive treatment, then women also receive treatment. Under the above constraints, the EWM rule chooses to give treatment to all groups except groups 1, 2, 3, 13, 14, and 15. The EWM rule determines not to treat unmarried Black or Hispanic individuals who worked for more than 12 weeks. On the contrary, the proposed shrinkage rule determines to treat such individuals when $\kappa=500$.
This study examined the problem of determining whether to treat individuals based on observed covariates. Particularly, we proposed a computationally tractable shrinkage rule that selects the shrinkage factor by minimizing the upper bound of the maximum regret. We also provided upper bounds of the ratio of the maximum regret of the shrinkage rule to those of the CES and pooling rules when the space of CATEs was correctly specified or misspecified. The theoretical and numerical results show that our shrinkage rule performs better than the CES and pooling rules in many cases when the space of CATEs is correctly specified. In addition, the results were robust to the misspecifications of the space of CATEs. Particularly, we found that the maximum regret of the shrinkage rule can be strictly smaller than that of the CES and pooling rules when the dispersion of the CATEs is moderate. Finally, we applied our method to experimental data from the JTPA study and showed that the decision of the shrinkage rule can differ from that of the CES rule even if $\kappa$ is a realistic value.
\setcounter{equation}{0}