EconBase
← Back to paper

Individual Shrinkage for Random Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

121,492 characters · 34 sections · 83 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Individual Shrinkage for Random Effects

\setstretch{1.1}

center[center omitted — 14 chars of source]
abstractThis paper develops a novel approach to random effects estimation and individual-level forecasting in micropanels, targeting individual accuracy rather than aggregate performance. The conventional shrinkage methods used in the literature, such as the James-Stein estimator and Empirical Bayes, target aggregate performance and can lead to inaccurate decisions at the individual level. We propose a class of shrinkage estimators with individual weights (IW) that leverage an individual's own past history, instead of the cross-sectional dimension. This approach overcomes the “tyranny of the majority" inherent in existing methods, while relying on weaker assumptions. A key contribution is addressing the challenge of obtaining feasible weights from short time-series data and under parameter heterogeneity. We discuss the theoretical optimality of IW and recommend using feasible weights determined through a Minimax Regret analysis in practice. Keywords: Micropanels; Shrinkage; Loss Function; Heterogeneity; Minimax Regret; Robustness JEL Classification: C10, C23, C53

\hfuzz=14pt

\setstretch{1.2} \parskip 2pt

Introduction

{\it “Knowing when to borrow and when not to borrow is one of the key aspects of statistical practice" (mallows1982overview)}

Estimating fixed or random effects (RE) and forecasting individual outcomes are core problems in econometrics, with application across different fields of economics. A significant econometric challenge arises when estimators rely on micropanels, as the short time dimension results in imprecise estimates.\footnote{Estimating individual effects and forecasting with micropanels are the goals of several literatures. One example is the vast literature using “value added models" to capture institutional effects, e.g., the effect of teachers on students' test scores Kane:Staiger:08,CFR:2014:a,CFR:2014:b, Angrist17, the effect of neighborhoods on intergenerational mobility CH18, or the effect of physicians on patients' outcomes FHB14. Hull20 discusses how such estimates play a key policy role in the regulation of healthcare and education in the U.S. Other literatures include kline2022systemic, who estimate firm-specific effects in order to analyze bias in firms' hiring decisions and chamberlain1999predictive, who consider forecasting individual incomes for consumption/savings decisions. Macroeconomic panel forecasting also falls into this category if it uses short estimation windows to account for parameter instability (e.g. Liu:2020 forecast banks' revenues after a regulatory change).} This challenge has motivated the widespread adoption of Bayesian shrinkage methods, such as the classical JS61's estimator (JS) and its extension by kwon21 or modern Empirical Bayes methods (EB), which “borrow strength" from other individuals to enhance accuracy. We argue that existing shrinkage methods, while improving aggregate performance, may lead to inaccurate decisions at the individual level. We propose a complementary shrinkage method designed to overcome the “tyranny of the majority" inherent in existing approaches.

Existing shrinkage methods minimize aggregate loss rather than individual loss. However, aggregate performance is not always the primary objective in many economic applications. For instance, when RE are estimated to guide policy interventions targeting specific individuals (e.g., teacher dismissal or hospital reward allocation, as in the value-added literature discussed by Hull20), forecast financial institution distress Liu:2020, or provide personalized financial advice chamberlain1999predictive, the focus is presumably on maximizing the accuracy of decisions for each individual, rather than optimizing aggregate performance.

Some pitfalls of targeting aggregate loss have been recognized in the Bayesian literature. For example, efron1971limiting and mallows1982overview highlight the “relevance" problem of the JS estimator, arguing that it suffers from the tyranny of the majority, by assuming that all other individuals are equally relevant for borrowing strength. Intuitively, JS shrinks individuals by the same amount regardless of their RE, which can result in large bias for outliers, i.e., individuals with RE that are far away from the common mean.\footnote{efron1971limiting suggests tackling the problem by first identifying and then not shrinking outliers; efron2010large proposes to use covariates to identify relevant individuals.} EB relies on stronger assumptions than JS, notably exchangeability and a distributional assumption for the errors. Exchangeability could be violated if a common RE distribution across individuals does not exist, raising concerns about robustness and external validity: RE estimates become dependent on the sample composition rather than reflecting true individual quality. Modern EB methods (e.g., efron2010large, gu2015unobserved,gu2023invidious, Liu:2020, chen2022empirical and koenker2024empirical) have made substantial progress in relaxing distributional assumptions on the prior; however, the impact of error distribution misspecification for EB remains largely unexplored. While there is some evidence that EB could be less susceptible to the tyranny of the majority under correct specification (see, e.g., some simulation results in Liu:2020), we show that a misspecified error distribution can result in large bias, particularly for outliers.

To address these limitations, we propose a shrinkage approach that explicitly targets individual loss. We introduce a class of shrinkage estimators that, similarly to JS, borrow strength from other individuals by shrinking the time series estimators of RE towards a common mean. Unlike existing approaches, however, our shrinkage with individual weights (IW) leverages solely the individual's own past history, rather than the cross-sectional dimension. Our goal is to be agnostic about parameter heterogeneity beyond the common mean, effectively representing the opposite end of the spectrum from existing literature, which assumes a shared parameter distribution across individuals.

While deriving optimal individual shrinkage rules (i.e., oracle weights) based on the Mean Squared Forecast Error (MSFE) criterion is straightforward in theory, the main challenge we address in this paper is constructing feasible (i.e., estimable) weights that perform well despite no restrictions on parameter heterogeneity and short time-series data. To tackle these issues, we employ the Minimax Regret criterion, which is well-suited when little is known about the parameter space. We first derive oracle weights that are minimax regret optimal relative to either the individual time series estimator or the pooled estimator. Although estimating these weights accurately in short samples remains challenging, we demonstrate that imposing bounds on a specific conditional expectation allows us to obtain feasible weights.

Our method does not rely on the strong assumptions inherent in EB methods. While EB necessitates the specification of a complete model with distributional assumptions on the errors, our method is semiparametric. Similarly to JS, IW can be viewed as a best linear rule that minimizes individual loss (whereas JS minimizes aggregate loss). IW does not assume a common RE distribution and thus accommodates richer unobservable heterogeneity, which we show is a crucial ingredient for obtaining accuracy at the individual level. In contrast, JS assumes homogeneous variances of the unobserved RE and errors. kwon21 can be interpreted as an extension of JS that assumes homogeneous variance for the RE but allows for heterogeneous variance of the errors. Standard applications of EB methods tend to assume a common RE distribution but can in principle allow for heterogeneous variances of the errors (e.g., gu2015unobserved and Section 5 of Liu:2020).

In alignment with our reference literature, we focus on a model where individual outcomes are expressed as the sum of RE and idiosyncratic errors. As previously mentioned, we accommodate unrestricted heterogeneity in parameters, while maintaining the assumption that RE share a common mean. This common mean is the point towards which we shrink the estimators/forecasts, representing how we borrow strength. The point of shrinkage is typically assumed to be zero if outcomes are demeaned, otherwise it is approximated by the pooled mean. Consistent with the literature, we also assume independence between RE and errors; however, we impose no restrictions on the relationship between the parameters that characterize their distributions. Notably, the outcomes in our model can also be interpreted as residuals from the first-step estimation of a linear panel data model or a value-added model with homogeneous coefficients for covariates, which may include lagged outcomes. Our framework thus encompasses a broad range of empirically relevant models.

In contrast to existing methods, an advantage of IW is that it does not depend on a large cross-sectional dimension for accuracy, which enables its application in small-sample settings. A drawback of relying on small samples is that it precludes the use of asymptotic behavior for evaluating the performance of IW. We focus instead on finite-sample optimality and on robustness, that is, on showing that IW performs well over the unknown parameter space. Our approach is inspired by manski2019econometrics, who emphasizes evaluation of decision rules by their performance across the parameter space and advocates the Minimax Regret criterion.\footnote{Minimax Regret properties of shrinkage estimators have been studied by magnus2002estimation and hansen2015shrinkage. However, their setting is different from ours in that they consider the problem of estimating the mean of normally distributed variables with known variance in the univariate case (magnus2002estimation) and in the multivariate case (hansen2015shrinkage). In addition, there is some literature applying the Minimax Regret criterion to panel data in order to address different questions: handling missing data in sample design Dominitz:Manski:2021 and forecasting discrete outcomes under partial identification or other concerns CMS:20. Their focus is also distinct from ours.} manski2019econometrics promotes this criterion as a robust, practical, and cautious decision-making framework, particularly effective under a high level of uncertainty. In our context, the Minimax Regret criterion not only provides a decision-theoretic foundation for the proposed class of shrinkage estimators, but it also aids in selecting the optimal weights in practice.

Since in our model the estimator of RE coincides with the forecast of the individual outcome, we can equivalently focus the discussion on estimation or forecasting. The theoretical results in the paper focus on forecasting, but we present the analogous results for estimation in the Appendix. We show that any IW forecast is Minimax Regret optimal over the parameter space, relative to using either the time series forecast or the common mean. In addition, IW is also optimal in terms of MSFE if we restrict attention to the region of the parameter space where the time series forecast and the common mean are equally accurate. Keeping all else equal, an additional improvement can be obtained under a key assumption that requires the IW weights to be genuine functions of the RE (in addition to not being “pathological", in the sense of shrinking outliers more than RE near the mean, which would exacerbate the tyranny of the majority phenomenon). JS, for example, does not satisfy this key assumption because its weights are based on cross-sectional information and thus are not functions of the RE. Finally, we show that the accuracy improvement of IW under the key assumption is larger the heavier the tails of the RE distribution. In other words, IW is particularly advantageous in the presence of heavy tails because the weights implicitly relate the amount shrinkage to how “far" the RE is from the common mean. This is what allows IW to overcome the tyranny of the majority.\footnote{The robustness of our approach is rooted in the focus on individual rather than aggregate accuracy. It is however worth noting that IW could dominate existing methods also in terms of aggregate accuracy. Whether this occurs in a given application depends on the distribution of the parameters across individuals - which our individual-level analysis leaves unrestricted - and also on the unknown tail properties of the RE distribution. The simulations and empirical applications offer some insights in this respect.}

We present three types of feasible weights for IW, all of which satisfy the key assumption: estimates of the oracle weights that are optimal in terms of the individual MSFE (IW-O); “Minimax Regret optimal weights" (IW-MR) and feasible weights based on the (in-sample or out-of-sample) inverse squared forecast error (IW-MSFE), which are equivalent to the weights considered in the forecast combination time-series literature (e.g., bates1969combination, stock1998comparison), in our case computed over a very short time series. These weights offer additional robustness benefits because they do not rely on correct specification of the model and can thus be applied in more general settings. We compare the finite sample performance of all feasible weights and conclude that IW-MR are the preferred weights, closely followed by (in-sample) inverse MSFE weights. Additional simulations illustrate how IW can overcome the tyranny of the majority phenomenon that affects JS.

We present two empirical illustrations. The first revisits the application in kline2022systemic, focusing on gender discrimination in firm hiring. We show that IW-MR delivers different results in terms of estimation and policy implications, compared to the EB procedure used in the paper. IW-MR also dominates in terms of forecasting performance and robustness/external validity. The second application forecasts individual earning residuals using the Panel Study of Income Dynamics. We find that the forecast with the best aggregate accuracy is IW-MR, which tends to assign high shrinkage weights to individuals with earning residuals near the median of the distribution. This application illustrates the potential usefulness of our approach even in terms of aggregate performance and even in highly heterogeneous environments where only a few individuals benefit from borrowing strength.

The rest of the paper is organized as follows. Section (ref) discusses the limitations of existing shrinkage approaches. Section (ref) summarizes our proposed IW method. Section (ref) shows the optimality of IW for the case of forecasting, both in terms of Minimax Regret and of MSFE, in a simplified setting with independent weights and forecasts. Section (ref) derives feasible weights for IW in a simplified setting with independent weights and forecasts. Section (ref) shows simulation results. Section (ref) discusses two empirical applications. Section (ref) offers concluding remarks. Appendix A contains the proofs, Appendix (ref) contains results regarding estimation and Appendix C formally derives the general feasible weights for IW presented in Section (ref).

Existing Methods and their Limitations

In this section, we illustrate how existing shrinkage methods that target aggregate loss - JS and EB - may deliver inaccurate individual-level decisions. We discuss the tyranny of the majority phenomenon as well as the effects of violating the exchangeability assumption or misspecifying the error term distribution.

Suppose that the individual outcomes are the sum of independent RE and errors:

align[align omitted — 88 chars of source]

where $A_i \sim (0, \lambda_i^2)$ and $U_{i,t} \sim (0, \sigma_i^2)$. Note the common zero mean and the heterogeneous variances. Denote by $\bar{Y}_i$ the $i$th-time series mean, i.e., the Maximum Likelihood Estimator:

align[align omitted — 66 chars of source]

We wish to estimate $A = \left(A_1, A_2, ..., A_N\right)$ or equivalently forecast $Y_{i, T+1}$ for each individual $i$ based on information up to time $T$, using a decision rule $\delta = \left(\delta_1, \delta_2, ..., \delta_N\right)$.\footnote{For notational simplicity we suppress the dependence on the sample.} Here we focus on estimation, but similar considerations apply to the forecasting problem.

An optimal rule can minimize individual or aggregate loss:

align*[align* omitted — 208 chars of source]

If one targets individual loss, restricting attention to linear rules delivers the optimal rule:\footnote{This can be obtained by applying the “best linear rule" in equation (9.4), page 129 of efron1973stein, to $\bar{Y}_{i}|A_i\sim(A_i,\sigma_i^2/T)$ and $A_i\sim(0,\lambda_i^2)$.}

align[align omitted — 114 chars of source]

The oracle IW rule thus shrinks the MLE towards the common zero mean, using individual-specific weights.

Targeting aggregate loss means using the posterior mean as the optimal rule. Here we briefly present the main Bayesian estimators that have been considered in the literature: the JS estimator of JS61 and EB estimators such as Liu:2020, efron2016empirical, and gu2015unobserved.

In the context of model ((ref)), JS does not require distributional assumptions but assumes homoskedasticity, $\sigma_i^2=\sigma^2$, $\lambda_i^2= \lambda^2$. JS61 and efron1973stein show that the following rule has lower aggregate risk, $\mathbb{E}_{A}L(\delta, A)$, than the MLE:

align[align omitted — 100 chars of source]

In practice, the JS weight is estimated by leveraging the cross-sectional dimension, as $\hat{\lambda}^2/(\hat{\lambda}^2+\hat{\sigma}^2/T)$, where $\hat{\sigma}^2/T=1/N \sum_{i=1}^N \left[1/(T-1)\sum_{t=2}^T (Y_{i,t}-Y_{i,t-1})^2\right]/(2T)$ and $\hat{\lambda}^2=1/N \sum_{i=1}^N(\bar{Y_i} - 1/N \sum_{i=1}^N\bar{Y_i})^2-\hat{\sigma}^2/T$. Notice that JS differs from the oracle IW rule in that it imposes constant weights across $i$, resulting in the same amount of shrinkage for all individuals.

EB methods require additional distributional assumptions on the error $U_{i,t}$ and estimate $A_i$ via the posterior mean. Two paradigms exist for computing the posterior mean: “G-modeling" and “f-modeling". In G-modeling, one first estimates the distribution of $A_i$, $G(A)$, leveraging the cross-sectional dimension, and then obtains the EB decision rule as the posterior mean based on the estimated $G(A)$. One approach to G-modeling is efron2016empirical, where the errors are assumed normal and homoskedastic and $G(A)$ is flexibly parameterized by a spline and estimated via deconvolution. Another G-modeling approach is gu2015unobserved, which uses a nonparametric maximum likelihood estimator based on the Kiefer-Wolfowitz approach for mixture models. gu2015unobserved assume normal errors that can be heteroskedastic (treating the variances as random). The G-modeling approach under normality and homoskedasticity results in the following rule:

align[align omitted — 207 chars of source]

with $\varphi _{i}\left( \cdot \right) $ the probability density function of a $\mathcal{N}\left(0,\sigma^{2}/T\right)$.

A different approach is f-modeling, which bypasses estimation of $G(A)$ and directly estimates the posterior mean using the Tweedie correction (available under the assumption that errors belong to the exponential family of distributions, see efron2011tweedie). The correction only depends on $G$ through the marginal density of the data, which can be estimated nonparametrically leveraging the cross-sectional dimension. Liu:2020 adopt this strategy, assuming normal and homoskedastic errors (Section 5 of Liu:2020 extends the method to allow for certain forms of heteroskedasticity in the errors), yielding the following rule:

align[align omitted — 106 chars of source]

where $l'(\bar{Y}_i)=\frac{\partial}{\partial \bar{Y}_i} \text{log} f(\bar{Y}_i)$ and $f(\bar{Y}_i)$ is the marginal density of $\bar{Y}_i$.

The Tyranny of the Majority

efron1971limiting, mallows1982overview and efron2010large discuss the notion of “relevance", which highlights a key challenge for Bayesian shrinkage methods: knowing which individuals are relevant for estimating a given $A_i$. efron2010large shows that ignoring relevance leads to bias and suggests using covariate information, when available, to define relevant individuals. To see why JS can lead to incorrect decisions at the individual level, assume a constant $\sigma^2_i=\sigma^2$ for simplicity. The individual-level bias incurred by using JS instead of the optimal rule (the oracle IW) is then: $$ Bias_i = \left(\frac{\lambda_i^2}{\lambda_i^2+\sigma^2/T} - \frac{\lambda^2}{\lambda^2+\sigma^2/T}\right)\bar{Y}_i.$$ As long as $\lambda_i^2 \neq \lambda^2$, the bias is large for large values of $\bar{Y}_i$, e.g., for outliers. This can be seen as an illustration of the tyranny of the majority phonomenon discussed by efron2010large.

There is some evidence that EB is less susceptible to the tyranny of the majority phenomenon under correct specification of the error distribution (see, e.g., simulation results reported in Figure 1 in Liu:2020).

Violation of Exchangeability

Related to the notion of relevance is the concept of exchangeability, which implies that the joint distribution of the data is invariant to permutations of the indices. Here we discuss the implications of the exchangeability assumption made by EB methods, specifically the assumptions of a common RE distribution across individuals combined with i.i.d. observations, which are sufficient to ensure exchangeability. Consider our fully heterogeneous model ((ref)) and maintain the normal errors assumption made by EB: $U_{i,t} \sim \mathcal{N}(0, \sigma_i^2)$. Then, the marginal likelihood is

align[align omitted — 373 chars of source]

where $L_i(\cdot | \lambda_i^2 )$ is the $i$-specific distribution of $A_i$ and $H_i(\cdot, \cdot)$ is the $i$-specific distribution of $(\lambda_i^2, \sigma_i^2)$. This likelihood function is too heterogeneous to be embedded in the EB approach, which thus requires additional assumptions. First, one typically assumes homogeneous $\lambda_i^2 \equiv \lambda^2$ and identical distributions $L_i$ and $H_i$ across $i$. Then, the marginal likelihood becomes

align[align omitted — 227 chars of source]

where $G( a, \sigma^2 ) := L (a) H (\sigma^2)$ is the unknown but common distribution. Second, EB methods typically assume that the random components $( A_i, \sigma_i^2)$ are i.i.d. over $i$, so that $G$ can be estimated nonparametrically. An EB decision rule is the posterior mean of the estimated $G$. Differences between the likelihoods in ((ref)) and ((ref)) imply that EB might lead to incorrect decision rules in our setting with heterogeneous $\lambda_i^2$ and no exchangeability assumptions.

Misspecification of the Error Distribution

EB methods rely on a complete model and assume a specific error distribution. The effects of misspecifying such distribution are largely unexplored in the literature. Here we illustrate the impact of misspecifying this distribution on EB estimators based on the Tweedie correction (e.g., Liu:2020) in a simple example.

For simplicity, assume $T=1$ and drop the $t$ subscript. Suppose that $U_{i}$ are distributed as a standardized Gamma distribution with zero mean, unit variance, and skewness $\gamma$.\footnote{Specifically, we consider $f_0(y)$ to be a standardized gamma variable with shape parameter m: $$f_0(y) \sim \frac{\operatorname{Gamma}_m - m}{ \sqrt{m}}$$ with $\operatorname{Gamma}_m $ having density $y^{m-1} \operatorname{exp}(-y)/m!$ for $y \geq 0$. Thus, $f_0(y)$ has mean 0, variance 1, and skewness $\gamma \equiv 2/\sqrt{m}$.} We can then obtain the misspecification bias of the EB estimator when wrongly assuming a standard normal distribution for the error term as:\footnote{Tweedie's formula under normal errors with zero mean and unit variance gives the posterior mean as: $$E[A_i|Y_i] = Y_i+ l'(Y_i),$$ where $l'(Y_i)=\frac{\partial}{\partial \bar{Y}_i} \text{log} f(Y_i)$ and $f(Y_i)$ is the marginal distribution of $Y_i$. Using efron2011tweedie, which shows Tweedie's formula when the error comes from a distribution in the exponential family, a Gamma distribution with shape parameter $m$ (and thus skewness $\gamma \equiv 2/\sqrt{m}$), zero mean and unit variance implies a posterior mean: $$E[A_i|Y_i] = \frac{Y_i+\gamma/2}{1+\gamma Y_i/2}+ l'(Y_i).$$ The bias in ((ref)) is obtained by subtracting the two expressions for $E[A_i|Y_i] $ under the assumption that $l'(Y_i)$ is the same in both cases as it is estimated from the data.}

align[align omitted — 87 chars of source]

We thus see that when $|Y_i|$ is large, e.g., for outliers, the misspecification bias is large in absolute value (the bias is approximately $\gamma/2$ when $Y_i$ is near the zero mean). This can be viewed as another manifestation of the tyranny of the majority phenomenon, here due to misspecification.

Shrinkage with Individual Weights (IW): Overview

Our focus on individual loss rather than aggregate loss allows us to overcome the issues described in Section (ref). In particular, we bypass the problem of relevance by only using the individual's past history to form the shrinkage weights instead of using the cross-sectional dimension, thus overcoming the tyranny of the majority. We do not require exchangeability or parametric assumptions on the error term and thus do not suffer from misspecification bias.

Our method is simple to implement. For each individual $i$, we propose shrinkage rules:

align[align omitted — 149 chars of source]

The point of shrinkage $\mu$ is either known (e.g. $\mu=0$ if the observations have been demeaned or if $Y_{i,t}$ are residuals from a first step estimation that includes an intercept) or is approximated with the pooled mean, $\mu= \Sigma^N_{i=1}\Sigma^T_{t=1}Y_{i,t}/NT$. $\widehat{Y}^{TS}_{i,T}$ is an estimator of the RE or a forecast of the outcome $Y_{i,T+1}$ that is only based on the time series dimension, typically the MLE: $\widehat{Y}^{TS}_{i,T}= \bar{Y}_{i,T}=\Sigma^T_{t=1}Y_{i,t}/T$. We derive three classes of feasible individual weights $W_{i,T}$, which we report here for convenience for the case where $\widehat{Y}^{TS}_{i,T}=\bar{Y}_{i,T}$.

Estimated Oracle Weights (IW-O):

align[align omitted — 264 chars of source]

where $(\cdot)^+$ denotes the positive part.

Minimax Regret Optimal Weights (IW-MR): These are the weights that perform best in our simulations:\footnote{A similar performance in simulations and in the empirical applications is obtained by the following IW-MR rule, based on an alternative unbiased estimator for $\sigma_i^2$:

align*[align* omitted — 169 chars of source]

}

align[align omitted — 177 chars of source]

Inverse MSFE Weights (IW-MSFE): These weights are based on the inverse MSFE (either in-sample or out-of-sample). Since they do not rely on the model assumptions, these weights are applicable in more general settings than the model considered in the paper. The in-sample inverse MSFE weights (which are the second best-performing weights in our simulations) are:

align[align omitted — 238 chars of source]

For a given choice of $P<T$, the out-of-sample inverse MSFE weights are:

align[align omitted — 256 chars of source]

where $\widehat{Y}^{TS}_{i,t-1}$ is the time series mean using data up to time $t-1$ (using all the data available or an arbitrary number of most recent data). For IW-MSFE-OOS, if $\mu$ is not known, it can be approximated with the pooled mean using data up to time $t-1$.

We note that the IW-MR feasible weights in ((ref)) ignore the possible dependence between $W_{i,T}$ and $\widehat{Y}^{TS}_{i,T}$. While we derive theoretical weights that account for such dependence in Appendix C, any attempt to estimate the additional term this introduces into the theoretical weights - a conditional expectation - would likely worsen the performance of the feasible weights. The feasible weights in ((ref)) thus approximate this additional conditional expectation term with the unconditional expectation, which is zero. Similar considerations motivate us to derive the results in Sections (ref) and (ref) in a simplified setting where $\widehat{Y}^{TS}_{i,T}=Y_{i,T}$ and the weights are based on information up to time $T-1$, which guarantees independence.

Optimality of IW

In this section we focus on forecasting and show conditions under which IW represents the optimal decision rule at the individual level in terms of two criteria: MSFE or Minimax Regret. In Appendix (ref), we show the analogous results for estimation.

Model and Assumptions

The model is:

align[align omitted — 73 chars of source]

where $A_i \sim (\mu, \lambda_i^2)$ and $U_{i,t} \sim (0, \sigma_i^2)$. Here, $A_i, U_{i,1},\ldots,U_{i,T}$ are random variables, whereas $\mu$, $\lambda_i^2$ and $\sigma_i^2$ are parameters. In other words, we take the frequentist approach.

This section considers a simplified setting for IW where the time series forecast is the time-$T$ outcome instead of the time series mean and the IW weights are based on data up to time $T-1$ instead of $T$. These assumptions imply that the weights and the time series forecasts are independent, which makes the theoretical results more transparent and intuitive. In this section we thus consider the following IW forecast:

align[align omitted — 234 chars of source]

where ${W}_{i,T-1}$ depends on time series data up to time $T-1$. We make the following assumptions.

assumption[Independence] $A_i, U_{i,1},\ldots,U_{i,T}$ are mutually independent.
assumption[Key Assumption] The individual weight ${W}_{i,T-1}$ satisfies $0 \leq {W}_{i,T-1} \leq 1$ and \begin{align} \mathbb{E} \left[ (A_i-\mu)^2 \left(1 - {W}_{i,T-1}\right)^2 \right] \leq \mathbb{E} \left[ (A_i-\mu)^2 \right] \mathbb{E} \left[ \left(1 - {W}_{i,T-1}\right)^2 \right], \end{align} or, equivalently, \begin{align*} \mathrm{Cov} \left\{ (A_i-\mu)^2, \left(1 - {W}_{i,T-1}\right)^2 \right\} \leq 0. \end{align*}
remark[Key Assumption] Assumption (ref) states that, as two random variables, $(A_i-\mu)^2$ and $\left(1 - {W}_{i,T-1}\right)^2 $ are weakly negatively correlated. Intuitively, the assumption requires that larger values of $(A_i-\mu)^2$ are associated with smaller weight attributed to the pooled forecast (or that the two are uncorrelated). This is a mild and reasonable assumption, in that it only rules out “pathological" weights that would shrink outliers more than individuals near the mean of the distribution, thus exacerbating the tyranny of the majority phenomenon that we are seeking to overcome. If the individual weight is a fixed constant, i.e., ${W}_{i,T-1} = c_i$ for some constant $0 \leq c_i \leq 1$, then Assumption (ref) is satisfied with an equality in (ref). Conversely, if the individual weight is a genuine function of the RE - instead of a fixed constant - the inequality in (ref) can be strict. We will see below that this strict inequality translates into improvements in the performance of IW.
remark[Interpretation of $\mu$] The common mean $\mu$ - the point of shrinkage - represents how we borrow strength from the majority. $\mu$ plays a similar role in our analysis as in a classical Bayesian setting, and we similarly consider it as a tuning parameter. As discussed by kwon21, in empirical work the outcomes are often first demeaned, in which case one simply sets $\mu=0$ (for example if $Y_{i,t}$ are residuals from a first-stage estimation of a model with an intercept, see remark (ref) below). If $\mu$ is unknown, we replace it with the sample mean of $Y_{i,t}$ over the panel. See remark (ref) below for a discussion of how $\mu$ could be chosen in case of a known group structure in parameters. Treating $\mu$ as a tuning parameter means that, similarly to existing approaches, our theoretical results do not take into account the uncertainty in its estimation.
remark[Extension: covariates] Covariates $X_{i,t}$ can be incorporated by redefining $Y_{i,t}$ in ((ref)) as residuals from the first-step estimation of a model with homogeneous coefficients: \begin{equation} Y_{i,t}=\tilde{Y}_{i,t}-X_{i,t}'\widehat{\beta}, \end{equation} where $\tilde{Y}_{i,t}$ are the outcomes and $\widehat{\beta}$ is a consistent estimator of the homogeneous coefficients as $N\rightarrow\infty$.\footnote{For instance, if the covariates $X_{i,t}$ include the lagged outcome, one could use the Arellano-Bond estimator arellano1991some to estimate $\beta$.} All the theoretical results discussed below then apply under the additional assumption that $N$ is large. Note that the assumption of consistency of $\widehat{\beta}$ could in principle be relaxed, as in a finite-$N$ setting there can be other, perhaps biased, estimators that improve forecast accuracy. We leave this extension for future research. The extension to a model with heterogeneous coefficients for the covariates would imply generalizing the problem of estimating/forecasting unobserved heterogeneity from the univariate case considered here to the multivariate case. While a full treatment of this extension is beyond the scope of this paper, we offer some remarks on this topic in the conclusion of the paper.
remark[Extension: value-added model] A value added model for an outcome $\tilde{Y}_{i,j,t}$ and covariates $X_{i,j,t}$ aims at estimating the RE $A_i$ in the model \begin{equation*} \tilde{Y}_{i,j,t}=X_{i,j,t}'\beta+A_i+U_{i,j,t}. \end{equation*} In the case of teacher value-added, for example, $i$ is the teacher and $j=1,...,n_{i,t}$ (with $n_{i,t}$ finite) are the students assigned to teacher $i$ in year $t$. This model can also be nested in ((ref)) if a consistent estimator (as $N\rightarrow\infty$) $\widehat{\beta}$ is available, in which case one defines $Y_{i,t}$ as residuals from a first step estimation involving averaged outcomes and covariates: \begin{equation} Y_{i,t}=\frac{1}{n_{i,t}}\sum_{j=1}^{n_{i,t}}\tilde{Y}_{i,j,t}-\frac{1}{n_{i,t}}\sum_{j=1}^{n_{i,t}}{X}_{i,j,t}'\widehat{\beta}. \end{equation}
remark[Robustness to distributional assumptions] Note that we make no distributional assumptions on the RE and the idiosyncratic errors. Heavy tails in both distributions are permitted, as long as the variances exist (i.e., the parameters $\lambda_i^2$ and $\sigma_i^2$ are finite).
remark[Robustness to dependence structure] The analysis below is carried out at the individual level, and thus in principle does not require a large $N$ nor any restriction on the cross-sectional dependence of $Y_{i,t}$. This is true as long as the point of shrinkage $\mu$ is known (e.g., when the outcomes have been demeaned so that $\mu=0$). When $\mu$ is unknown and approximated with the sample mean over the panel, this implicitly requires a restriction on the cross-sectional dependence that ensures validity of a law of large numbers. The incorporation of covariates discussed in remark (ref) also implicitly restricts the dependence structure by assuming availability of a consistent estimator of the homogeneous coefficients. Time-series dependence can be accounted for by including lagged dependent variables as covariates, as long as the autoregressive coefficients are homogeneous across individuals.
remark[Robustness to distribution of parameters across individuals] Our analysis is carried out at the individual level and is purposely agnostic about the distribution of $\lambda_i^2$ and $\sigma_i^2$ across $i$. This implies that, in general, we cannot make any formal statement about the aggregate performance of our estimator. Nonetheless, we are able to provide some intuition for the implications of our findings for aggregate accuracy, see Section (ref) below. Another implication is that, while we assume independence between individual RE and errors, we accommodate any type of unknown relationship between their variances (e.g., there could be two groups of individuals, one with low $\lambda_i^2$ and low (high) $\sigma_i^2$ and one with high $\lambda_i^2$ and high (low) $\sigma_i^2$). The next remark highlights how the analysis can be modified if one is willing to assume a known group structure in parameters.
remark[Known group structure in parameters] Suppose there is a group structure in $\mu$, with a finite number of subgroups and observable group membership (but with $\lambda^2_i$ and $\sigma^2_i$ still heterogeneous within the subgroups). In this case, the only modification to our analysis is that the point of shrinkage is the mean for the subgroup instead of the mean for the whole panel. If the homogeneity assumption within subgroups extends to $\lambda^2_i$ and $\sigma^2_i$, then our estimator becomes the JS estimator applied to each subgroup (and thus it is exactly the JS estimator if there is only one group).

Henceforth, we focus on model ((ref)), with the understanding that $Y_{i,t}$ are either raw outcomes or residuals such as ($\ref{panelres})$ or ($\ref{vamres}$) (in a large-$N$ setting).

MSFE and Minimax Regret

This section discusses the two criteria that we consider for evaluating the performance of IW: MSFE and Minimax Regret. Consider a situation where there is uncertainty about the parameter $\theta_i = (\lambda^2_i,\sigma^2_i)$. The MSFE of forecast $m \in \mathcal{M}$ for a given $\theta_i$ is

\[ \mathrm{MSFE}(m, \theta_i)=\mathbb{E} \left[ \left( Y_{i,T+1} - \widehat{Y}^{m}_{i,T} \right)^2 \right]. \]

The next lemma derives the MSFEs of TS, Pool and IW in ((ref)).

lemmaConsider the forecasts in ((ref)). Then under Assumption (ref) we have \begin{align*} \mathrm{MSFE}(\mathrm{TS}, \theta_i) &= 2 \sigma_i^2, \\ \mathrm{MSFE}(\mathrm{Pool}, \theta_i) &= \lambda_i^2 + \sigma_i^2, \\ \mathrm{MSFE}(\mathrm{IW}, \theta_i) &= \sigma_i^2 + \sigma_i^2 \mathbb{E} \left[ {W}_{i,T-1}^2 \right] + \mathbb{E} \left[ (A_i - \mu)^2 \left(1 - {W}_{i,T-1}\right)^2 \right]. \end{align*}

Lemma (ref) suggests that the trade-off between TS and Pool in terms of MSFE depends on the “signal-to-noise" ratio $\lambda_i^2/\sigma_i^2$: Pool dominates when the ratio is less than 1 and TS dominates when it is greater than 1. Knowing the parameters would allow one to choose the best forecast for individual $i$; however in the presence of uncertainty about the parameters it is not possible to choose a forecast optimally. We thus pursue an alternative route. We seek a robust rule that performs well over the entire parameter space, in the sense of avoiding large errors when TS and Pool have different accuracy and improving on both TS and Pool when they have similar accuracy. The following sections show that IW can accomplish both goals.

We first formalize the notion of robustness that we consider here, based on the Minimax Regret criterion. Let $\mathcal{M}$ include $\mathrm{TS}$, $\mathrm{Pool}$, and $\mathrm{IW}$. We define regret as

align[align omitted — 131 chars of source]

The Minimax Regret (MMR) criterion selects the forecast $m$ that minimizes the maximum regret \[ \max_{\theta_i \in \Theta} R(m,\theta_i), \] where $\Theta$ is the parameter space. The form of regret here is similar to that of regret in decision theory without sample data (e.g., see equation (3) in manski2019econometrics). The MMR criterion is championed by manski2019econometrics.\footnote{See Section A.2 in manski2019econometrics and references therein for a detailed discussion.} The regret in (ref) is defined relative to the best forecast (in terms of MSFE) out of a set of three because the goal in this section is to choose among IW, TS, and Pool.\footnote{If we had adopted the criterion of Minimax instead of Minimax Regret, we would have an obvious but trivial solution that TS is preferred to Pool if $\max_{i} \sigma_i^2 < \max_{i} \lambda_i^2$ and vice versa. In this setting, it is not necessarily the case that IW provides a better performance in terms of minimizing the maximum MSFE.}

Minimax Regret Optimality of IW

In this section, we show the conditions under which IW is optimal in terms of Minimax Regret. We restrict our attention to the parameter space represented in Figure 1 below, where the signal-to-noise ratio $\lambda_i^2/\sigma_i^2$ ranges from $1-\nu$ to $1+\nu$ for some $0 \leq \nu < 1$.

align[align omitted — 155 chars of source]
figure[figure omitted — 177 chars of source]

Considering a neighbourhood of 1 is natural, as we saw that this point represents the case where TS and Pool are equally accurate. The radius of the neighbourhood is constrained by the fact that the signal-to-noise ratio cannot be negative, so in practice we only exclude cases where TS strongly dominates, due to large variance of the RE and low variance of the error.\footnote{The adoption of a common $\nu$ for the upper and lower bounds is only for convenience in deriving analytical results. Figures (ref) and (ref) below make it clear that we are being conservative, as increasing the upper bound on $\lambda_i^2/\sigma_i^2$ would not change the conclusions of the Minimax Regret analysis.}

The next theorem shows that IW (uniquely) minimizes maximum regret among TS, Pool, and IW under Assumptions (ref) and (ref).

theoremLet Assumptions (ref) and (ref) hold. Then, \begin{align*} \max_{\theta_i \in \Theta} R(\mathrm{IW},\theta_i) \leq \min \left\{ \max_{\theta_i \in \Theta} R(\mathrm{TS},\theta_i), \max_{\theta_i \in \Theta} R(\mathrm{Pool},\theta_i) \right\}, \end{align*} where $\Theta$ is defined in (ref). Furthermore, the inequality above is strict if either $0 < {W}_{i,T-1} < 1$ with positive probability or the inequality in (ref) is strict.

The theorem shows when the improvement of IW over TS and Pool in terms of regret is strict. For example, any constant weight between 0 and 1 provides an improvement. Furthermore, keeping all else equal, a weight that is a genuine function of the RE that satisfies Assumption (ref) with a strict inequality will deliver an additional improvement in performance. Existing Bayesian shrinkage approaches such as the JS estimator, for example, deliver weights that are strictly between 0 and 1 but that do not depend on $A_i$. This means that JS outperforms (in terms of Minimax Regret) using TS or Pool for all individuals, but it can be, in turn, outperformed by any weight satisfying Assumption (ref) that is a genuine function of the RE. The theorem thus illustrates the potential benefits of individual weights that are based on information in the time series dimension, and thus capture the RE, relative to existing shrinkage approaches that leverage instead the cross-sectional dimension.

We illustrate the findings of Theorem (ref) in Figures (ref) and (ref). Consider one individual (so drop the subscript $i$) observed over 4 time periods, with $U_{1},...,U_{4}$ drawn independently from a $\mathcal{N}(0, 1)$ and $A$ drawn from a $\mathcal{N}(0, \lambda^2)$. Repeating the simulation a large number of times allows us to approximate the individual MSFE and regret when forecasting $Y_4$ at time $T=3$ using either TS, Pool or IW. The figures plot these MSFEs and regrets as a function of the signal-to-noise ratio. For IW, we consider the feasible Minimax Regret optimal rule (IW-MR) that we derive in equations ((ref)) and ((ref)) below.\footnote{ Specifically, we have $\widehat{Y}^{TS}_{3}=Y_3, \widehat{Y}^{Pool}_{3}=0$ and $\widehat{Y}^{IW-MR}_{3}=Y_{3} {W}_{2}$, with $W_2= 1 - 1/\sqrt{\frac{ \max \{ Y_{1}^2,Y_{2}^2 \}}{0.5(Y_{1} - Y_{2})^2 } + 1}$.} Figure (ref) shows that no forecast uniformly dominates in terms of MSFE over the parameter space; however, IW is the most accurate over the majority of the parameter space, except for very small values of the signal-to-noise ratio, when Pool dominates. Figure (ref) shows that IW is additionally Minimax Regret optimal. To see why, note that the regret for TS (dashed line) achieves a maximum value of around 1 when the signal-to-noise ratio is close to zero. The regret for Pool (dotted line) obtains a maximum value of around 1.4 when the signal-to-noise ratio is large. The regret for IW (solid line) achieves a maximum value of around 0.27, when the signal-to-noise ratio is close to zero. Thus, over the parameter space, the Minimax Regret optimal rule is IW, since it has the smallest maximum regret among the three rules.

figure[figure omitted — 164 chars of source]
figure[figure omitted — 167 chars of source]

MSFE Optimality of IW

Figure (ref) suggests that IW outperforms TS and Pool in terms of MSFE when the signal-to-noise ratio is in a neighborhood of 1. One may wonder if this result holds regardless of the data-generating process. The following theorem shows that, indeed, IW outperforms TS and Pool in terms of MSFE regardless of the data-generating process, if we restrict attention to the case where the signal-to-noise ratio equals 1. This means that IW is not only robust - i.e. optimal in terms of Minimax Regret - but it is also optimal in terms of MSFE when TS and Pool are equally accurate and thus would be indistinguishable.

theoremLet Assumptions (ref) and (ref) hold. Suppose that $\lambda_i^2 = \sigma_i^2$. Then, \begin{align*} \mathrm{MSFE}(\mathrm{IW}, \theta_i) &\leq \mathrm{MSFE}(\mathrm{TS}, \theta_i) = \mathrm{MSFE}(\mathrm{Pool}, \theta_i) = 2 \sigma_i^2. \end{align*} Furthermore, the inequality above is strict if either $0 < {W}_{i,T-1} < 1$ with positive probability or the inequality in (ref) is strict.

The theorem shows that IW is weakly more accurate than TS and Pool when the two forecasts have equal accuracy. Furthermore, as in the case of the Minimax Regret optimality results in Theorem (ref), a strict accuracy improvement can be obtained when the weights are strictly between 0 and 1 or are genuine functions of the RE. This shows that considering individual weights that leverage time series information to capture the RE can deliver a strict improvement in accuracy in terms of MSFE.

Accuracy Gains and Tail Heaviness

In this section we perform a simulation exercise to illustrate the result of Theorem (ref) and further show how the accuracy gains of IW are linked to the heaviness in the tails of the RE distribution. As for Figures (ref) and (ref), we consider one individual observed over 4 time periods. We now however focus on the case $\sigma^2=\lambda^2=1$ (making TS and Pool equally accurate) with $U_{1},...,U_{4}$ drawn independently from a $\mathcal{N}(0, 1)$ and $A$ drawn from a Pareto distribution with different degrees of tail heaviness (in a way that ensures that $\sigma^2=\lambda^2$).\footnote{We consider the double Pareto distribution with pdf $f(x; \theta; \beta)=\theta/(2\beta)

cases(x/\beta)^{\theta-1}, & if $0<x<\beta$ \\ (\beta/x)^{1-\theta}, & if $x \geq \beta$

$ with the following parameter combinations for the shape ($\theta$) and scale ($\beta$) parameters: $(2.3, .5)$, $(3,1)$, $(5, 2.45)$, $(50, 34.5)$. Note that population moments of order $\theta$ or greater do not exist. We thus quantify the tail heaviness of the distribution of $A_i$ by reporting a robust quantile-based measure of kurtosis, the Crow-Siddiqui measure ($CS= (Q_{0.975}-Q_{0.025})/(Q_{0.75}-Q_{0.25})$). Tor each of the above four combinations, this respectively equals 7.58, 6.59, 5.51, 4.42.} Repeating the simulation a large number of times allows us to approximate the individual MSFE when forecasting $Y_4$ at time $T=3$ using either TS, Pool or IW. For IW we consider the feasible rule IW-MR in footnote \ref{footIWMR}. Figure \ref{fig:mmr_equallyaccurate} plots the MSFEs of TS, Pool and IW as a function of the heaviness in the tails of the distribution of $A$, as captured by the Crow-Siddiqui measure of kurtosis (on the x-axis). The figure shows that IW improves on the performance of TS and Pool when TS and Pool are equally accurate, confirming the findings of Theorem (ref). Furthermore, the accuracy gains of IW are larger the heavier the tails of the RE.

figure[figure omitted — 231 chars of source]

Validity of Assumption (ref) and Relationship with Tail Heaviness

It is easy to verify that all the feasible weights reported in Section (ref) satisfy Assumption (ref) with a strict inequality. For all weights, the term $(A_i-\mu)^2$ appears in the denominator of $1-W_{i,T}$, whereas the remaining terms only depend on the $U_{i,t}'s$. This implies that large values of $(A_i-\mu)^2$ are associated with small values of the weight on the pooled forecast $\mu$. To illustrate how Assumption (ref) is linked to tail heaviness, we consider the same simulation design as that obtained to produce Figure (ref) and compute the covariance in Assumption (ref) (focusing for simplicity on IW-MR only). The four distributions have increasing tail heaviness, while everything else that would otherwise affect the weights is kept fixed. We find that this covariance becomes more negative as the tail heaviness increases (it respectively equals -0.134, -0.168, -0.216, -0.254 for the four levels of tail heaviness). This implies that heavy tails make the inequality in Assumption (ref) more pronounced, which, as shown by Theorems (ref) and (ref), translates into larger gains of IW relative to TS and Pool.

Implications for Aggregate Performance

The findings in the previous sections show the benefits of IW in terms of individual performance, but they also have implications for aggregate performance. Figure (ref), in particular, provides some intuition for how the distribution of the signal-to-noise ratio $\lambda^2_i/\sigma^2_i$ across $i$ (which we do not restrict in any way) affects aggregate performance as measured by the average MSFE. If there are enough individuals for which the signal-to-noise ratio is in the range where IW dominates in terms of individual MSFE, IW will dominate also in terms of average MSFE. In addition, our simulations below will show how the tail properties of the distribution of RE can be linked to improved performance of IW relative to shrinkage estimators. Another simulation will show that IW can be beneficial not only for outliers, but also for individuals that are near the mean of the distribution. This implies that improvements in terms of aggregate performance of IW relative to existing methods can be linked to how many individuals fall in the tails and/or near the mean of the RE distribution.

Feasible Weights for IW

The results in the previous section show that individual weights are optimal under assumption (ref), but do not directly provide a way to derive feasible weights. In this section, we show that we can derive feasible weights that satisfy this assumption. Here we describe three types of weights. For the first two, as in the previous section, we focus on the simplified setting in section (ref) for illustrative ease, but, again, the weights that we propose to use in practice and that we consider in the simulations and empirical application are the general weights reported in section (ref) and derived in Appendix C under general conditions. The third type of weights are not based on the model assumptions and thus are the same as in the general case.

Estimated Oracle Weights (IW-O)

The first set of feasible weights are based on the oracle weights that minimize the individual MSFE, $$\mathrm{MSFE}(\widehat{Y}^{IW}_{i,T} )=\mathbb{E} \left[ \left( Y_{i,T+1} - \widehat{Y}^{IW}_{i,T} \right)^2 \right],$$ which are functions of the individual variance parameters\footnote{These oracle weights follow for example from equation (9) in Chapter 4 of timmermann2006forecast, using the fact that the joint distribution of $Y_{i,T+1}$ and ${Y}_{i,T}$ is $$

pmatrix[pmatrix omitted — 40 chars of source]

\sim \left(

pmatrix[pmatrix omitted — 25 chars of source]

,

pmatrix[pmatrix omitted — 95 chars of source]

\right),$$ which gives the optimal weight on ${Y}_{i,T}$ as the product between the inverse of the variance of the forecast and the covariance between the outcome and the forecast. The linear combination with $W^{o}_{i}$ as weight could also be obtained by applying the ``best linear rule" in equation (9.4), page 129 of \cite{efron1973stein}, to ${Y}_{i,T}|A_i\sim(A_i,\sigma_i^2)$ and $A_i\sim(0,\lambda_i^2)$.}:

align[align omitted — 81 chars of source]

Estimated oracle weights at time $T-1$ can be obtained as

align[align omitted — 195 chars of source]

These weights use the fact that: $\Sigma^{T-1}_{t=1} (Y_{i,t}-\mu)^2 /(T-1)$ is an unbiased estimator of $\lambda_i^2+\sigma_i^2$ and that $\widehat{\sigma}^2_i=\Sigma^{T-2}_{t=1} (Y_{i,t} - Y_{i,t+1})^2 /2(T-2)$ is an unbiased estimator of $\sigma_i^2$.\footnote{To see that $\widehat{\sigma}^2_i$ is an unbiased estimator of $\sigma_i^2$ note that:

align*[align* omitted — 197 chars of source]

}

The short time dimension leads to imprecise estimates of these parameters, which can negatively impact the performance of feasible oracle weights. Furthermore, the subtraction in the numerator may produce negative weights. While taking the positive part - as we do for the general weights reported in Section (ref) and derived in Appendix C - partially alleviates this issue, our simulations indicate that these weights still perform poorly in practice.

These considerations motivate our focus on developing feasible weights that are robust, specifically those that optimize Minimax Regret across the unknown parameter space.

Minimax Regret Optimal Weights (IW-MR)

In order to obtain feasible Minimax Regret optimal weights we shift from unconditional MSFE to MSFE that is conditional on the information set at time $T-1$. The following lemma is the analog of Lemma (ref) for the conditional MSFE.

lemmaConsider the forecasts in ((ref)). Let Assumption (ref) hold. Then, the MSFEs conditional on the information set at time $T-1$, $\{Y_{i,1},...,Y_{i,T-1}\}$ are \begin{align*} \mathrm{MSFE}(\mathrm{TS}, \theta_i | \{Y_{i,1},...,Y_{i,T-1}\}) &= 2 \sigma_i^2, \\ \mathrm{MSFE}(\mathrm{Pool}, \theta_i | Y_{i,1},...,Y_{i,T-1}) &= \kappa_{i,T-1}^2 + \sigma_i^2, \\ \mathrm{MSFE}(\mathrm{IW}, \theta_i | Y_{i,1},...,Y_{i,T-1}) &= \sigma_i^2 (1 + {W}_{i,T-1}^2 ) +\kappa_{i,T-1}^2 \left(1 - {W}_{i,T-1}\right)^2, \end{align*} where \begin{align} \kappa_{i,T-1}^2 := \mathbb{E} \left[ (A_i-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right]. \end{align}

It is easy to verify that the weights that minimize $\mathrm{MSFE}(\mathrm{IW}, \theta_i | Y_{i,1},...,Y_{i,T-1})$ are given by:

align[align omitted — 89 chars of source]

In this section, we consider a different type of regret, defined as the difference between the conditional MSFE for a generic weight $W_{i,T-1}$ and the conditional MSFE corresponding to the conditionally optimal weights $W^*_{i,T-1}$ in (ref):

align[align omitted — 530 chars of source]

where

align[align omitted — 172 chars of source]

The form of regret in (ref) is similar to that of regret in statistical decision theory (e.g., see equation (6) in manski2019econometrics).

The following theorem derives the optimal minimax regret weights under the assumption that we can bound the random variable $\zeta_{i,T-1}^2$.

theoremLet Assumption (ref) hold. Suppose that $\zeta_{i,T-1}^2$ in ((ref)) is such that $\zeta_{i,T-1}^2 \in [0, \tilde{\zeta}_{i,T-1}^2]$, where $\tilde{\zeta}_{i,T-1}^2$ is large enough that maximum regret occurs at $\zeta_{i,T-1}^2 = \tilde{\zeta}_{i,T-1}^2$ with positive probability. Consider maximum regret \begin{align*} &\max_{\theta_i \in \Theta} R^*(W_{i,T-1}, \theta_i | Y_{i,1},...,Y_{i,T-1}) \\ &= \sigma_i^2 \max \left[ {W}_{i,T-1}^2, \left\{ {W}_{i,T-1}^2 + \tilde{\zeta}_{i,T-1}^2 \left(1 - {W}_{i,T-1}\right)^2 - \frac{\tilde{\zeta}_{i,T-1}^2}{\tilde{\zeta}_{i,T-1}^2 + 1} \right\} \right], \end{align*} with $R^*(W_{i,T-1}, \theta_i | Y_{i,1},...,Y_{i,T-1})$ defined as in ((ref)). Then, the weight that minimizes maximum regret is \begin{align} {W}^{IW-MR}_{i,T-1} = 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T-1}^2 + 1}}. \end{align}

In practice, the value of the bound $\tilde{\zeta}_{i,T-1}^2$ is uncertain, but the following heuristic rule can be used to obtain feasible weights. Assuming $T \geq 3$, we define

align[align omitted — 196 chars of source]

where $\mu$ is either known or approximated by the pooled mean. Intuitively, the denominator \[ \sum_{t=1}^{T-2} (Y_{i,t} - Y_{i,t+1})^2 / [2(T-2)] \] is an unbiased estimator of $\sigma_i^2$, the denominator of $\zeta_{i,T-1}^2$. The numerator serves as a proxy for an upper bound on $\kappa_{i,T-1}^2$, the numerator of $\zeta_{i,T-1}^2$.

Although the construction is heuristic, we may interpret our result as a minimax-regret optimal rule conditional on $\tilde{\zeta}_{i,T-1}^2 = \widehat{\tilde{\zeta}_{i,T-1}^2}$. This is similar in spirit to partial identification settings where the outcome variable is known to lie within a bounded interval (e.g., $Y \in [y_{\min}, y_{\max}]$). When $y_{\min}$ and $y_{\max}$ are unknown, it is common to use the sample minimum and maximum as proxies, and interpret the resulting identification region as conditional on these sample bounds.

Inverse MSFE Weights (IW-MSFE)

The weights we derive in this section do not rely on the model and the assumptions, and are thus applicable in more general settings. The weights are based on comparing the (in-sample or out-of-sample) MSFE at time $T$ of the TS and Pool forecasts. These weights are analogous to those considered in the time-series forecast combination literature (e.g., bates1969combination, stock1998comparison), with the difference that the MSFE is computed here for each individual over a very small time series sample (possibly containing only one observation). As in stock1998comparison, the weights ignore any correlation between the TS and the Pool forecasts.\footnote{In the time series literature, these weights are known to perform well even when the time dimension is large because of the challenges in estimating correlations precisely. See, e.g., the discussion in stock1998comparison.}

The in-sample inverse MSFE weights are given by:

align[align omitted — 260 chars of source]

The out-of-sample inverse MSFE weights are given by:

align[align omitted — 183 chars of source]

Note that here we base ${W}_{i,T}^{OOS} $ only on the out-of-sample forecast errors at time $T$ corresponding to the TS and Pool forecasts computed on the sample up to time $T-1$. Depending on the magnitude of $T$, one could also compute the out-of-sample MSFEs using more than just one out-of-sample period. For example, one could select $P<T$ and consider

align[align omitted — 280 chars of source]

Finally, we note that one could consider “rolling-window" forecasts, both as the original TS and Pool forecasts and in the computation of the weights. In this case, both TS and Pool forecasts at time $t$ would be based only on the $R<t$ most recent observations, rather than all available observations up to time $t$.

Monte Carlo Simulations

In this section we first study the finite sample performance of alternative feasible IW weights. We then compare IW to JS.

Comparing Feasible IW rules

We consider one individual (so we here drop the subscript $i$) observed over 3 time periods, with errors $U_{1},...,U_{3}$ drawn independently from a $\mathcal{N}(0, 1)$ and RE $A$ drawn from a $\mathcal{N}(0,\lambda^2)$, with $\lambda^2$ taking $50$ equally-spaced values on the grid $[0.001, 2]$. Repeating the simulation 10000 times allows us to approximate the individual MSFE when forecasting $Y_3$ at time $T=2$ using the different IW feasible weights reported in Section (ref) (with $\widehat{Y}_{i,T}^{TS}=\bar{Y}_i$ and $\mu=0$). Figures (ref) and (ref) respectively report the MSFEs of each IW rule, divided by the MSFE of IW-MR, and the Regret of each IW rule as a function of the signal-to-noise ratio (which here equals $\lambda^2$). In Figure (ref), a line above 1 means that the rule is dominated by IW-MR.

figure[figure omitted — 195 chars of source]
figure[figure omitted — 199 chars of source]

Figures (ref) and (ref) illustrate the dominance of IW-MR (black solid line) over the other feasible rules, in terms of both MSFE and Regret. Figure (ref) shows that IW-MR is Minimax Regret optimal over the parameter space. IW-MSFE-IS (blue dotted line) is uniformly dominated in terms of both criteria by IW-MR, although by a small amount. The performance of IW-MSFE-OOS (red dashed line) and IW-O (green dashed-dotted line) depends on the signal-to-noise ratio, both outperforming IW-MR when the signal-to-noise ratio is very low, but performing poorly over the rest of the parameter space. We thus conclude that IW-MR is the preferred rule, closely followed by IW-MSFE-IS.

IW vs. JS

In this section we compare the performance of IW-MR to that of the JS forecast (with estimated weights, as reported after equation ((ref))). The draws in the following simulations can be interpreted in two different ways. First, they can be seen as different possible draws of the RE for one individual. Second, they can be seen as draws for different individuals that have the same distribution of the RE. Averages across simulations accordingly have a different interpretation: they approximate the individual MSFE in the first interpretation and the aggregate MSFE in the second. The JS forecast in the first interpretation is simply an IW with constant weights that do not depend on the RE (see the discussion after Theorem (ref)). In the second interpretation, JS is the forecast that exploits information from the cross-sectional dimension, in contrast to IW, which leverages the time series dimension. Note that the assumptions of parameter homogeneity made by JS are satisfied in all the designs we consider.

Tyranny of the Majority

We start by visually illustrating how IW overcomes the tyranny of the majority phenomenon that affects JS. Henceforth, we focus on the IW-MR rule, which we saw in the previous section generally outperforms the other feasible rules. We consider 10000 simulations of outcomes generated as $Y_{t} = A + U_{t}$, with $t=1,...,3$, $U_{t} \sim \mathcal{N}(0,1)$, independent across $t$. For the RE we consider the following designs:

itemize• Design 1 (Normal): $A \sim \mathcal{N}(0,\lambda^2)$, where $\lambda^2 \in \{1,3\}$. • Design 2 (Laplace): $A \sim Laplace$ with parameters (0,1), which implies mean 0 and variance $\lambda^2=2$. • Design 3 (Double Pareto): $A \sim Double Pareto(\theta,\beta)$, where $\theta=3$ and $\beta=1$ (which implies mean 0 and variance $\lambda^2$ around $1.1$).

These designs correspond to an increasing heaviness in the tails of the RE distribution.

We compare IW-MR as described in Section (ref) (with $\widehat{Y}^{Pool}_T=0$) to JS (with estimated weights, as reported after equation ((ref))).

Figures (ref) - (ref) report the difference $\Delta SFE$ between the squared forecast errors of forecasts made at $T=2$ for IW-MR versus JS, for the different designs. The horizontal axis reports the value of $A$. The figures illustrate the tyranny of the majority phenomenon: JS tends to make larger errors than IW-MR (the dots fall below zero) for RE in the tails and also near the center of the distribution. This pattern is not yet visible in Figure (ref) for the normal design with low variance, where the cloud appears symmetric relative to the horizontal axis, but it is clear in the remaining figures. For example, in Figure (ref) (the normal design with larger variance) the cloud is heart-shaped, showing the superior performance of IW-MR near the center of the distribution. Figure (ref) (the Laplace design) also shows the heart shape but also the superior performance of IW in the tails. The improvement in the tails is starkly evident in Figure (ref) (the Double Pareto design), where the cloud has an inverted U-shape. These results illustrate that what matters for the tyranny of the majority is not only the tail heaviness of the RE distribution, but its relationship to the variance: the design in Figure (ref) shows that the phenomenon is present even when the distribution has thin tails, but large variance. This is intuitive, as both a high variance and heavy tails make it worthwhile to link the shrinkage to the RE (IW) instead of shrinking every individual by the same amount (JS).

figure[figure omitted — 179 chars of source]
figure[figure omitted — 175 chars of source]
figure[figure omitted — 176 chars of source]
figure[figure omitted — 182 chars of source]

Aggregate Performance

The second interpretation of the simulation discussed above, which views the different draws as different individuals, allows us to also analyze aggregate performance. Averaging the $\Delta SFE$'s reported in each figure across $i$ gives a measure of the relative aggregate performance of IW-MR and JS. These average equal 0.019, 0.025, -0.005 and -0.027, respectively in Figures (ref) - (ref). The relative aggregate performance thus depends on the tail properties of the distribution of RE, with JS dominating in the normal cases and IW-MR dominating in the heavy-tailed cases.

Empirical Applications

We consider two applications of IW (specifically, IW-MR from Section (ref)).

Estimating and Forecasting Systemic Firm Discrimination

In this section we use IW to extend the analysis in kline2022systemic, assessing the extent to which large U.S. employers systemically discriminate job applicants based on gender. We compare the performance of IW-MR to that of JS (with estimated weights, as reported after equation ((ref))) and EB (specifically, the deconvolution estimator of efron2016empirical, henceforth Efron).

Data

We use the panel dataset in kline2022systemic on an experiment that consisted of sending fictitious applications to jobs posted by 108 of the largest U.S. employers. For each firm, 125 entry-level vacancies were sampled and, for each vacancy, 8 job applications with random characteristics were sent to the employer. Sampling was organized in 5 waves (between October 2019 and April 2021). Focusing on firms sampled in all waves yields a balanced panel of $N=72$ firms over $T=5$ waves.\footnote{Accounting for vacancy closures and the exclusion of some firms from some waves reduced the number of applications to $65,400$.}

Applications were sent in pairs, one randomly assigned a distinctively female name and the other a distinctively male name. For details on the other observables see kline2022systemic. The primary outcome in kline2022systemic is whether the employer attempted to contact the applicant within 30 days of applying. The gender contact gap is defined as the firm-level difference between the contact rate (the ratio of number of contacts and number of received applications) for male and that for female applications.

EB approach

The results in kline2022systemic are based on Efron. The approach considers firm-specific studentized contact gaps, $y_{i.t}=Y_{i,t}/s_i$, where $Y_{i,t}$ is the contact gap and $s_i$ is the standard deviation of contact gaps across different job applications for firm $i$. These are modelled as

$$y_{i,t}=a_i+u_{i,t}, \;\;\;\;\; u_{i,t}\sim \mathcal{N}(0,1) \;\;\;\;\; a_i \sim G_{a}, \;\;\;\;\; \text{for} \; i=1, ..., 72. $$ $G_{a}$ is assumed to belong to an exponential family, parameterized by a fifth-order spline. By pooling observations from all five waves, Efron yields penalized Maximum Likelihood estimates of the spline parameters and thus an implied distribution $\hat{G}_{a}$ of studentized contact gaps with density $\hat{g}_{a}= d \hat{G}_{a}$. One can then recover the distribution $\hat{G}_{A}$ of the RE for the unstudentized contact gaps $Y_{it}$ under the assumption of independence between the RE and $s_i$. In particular, the density $\hat{g}_{A} = d \hat{G}_{A}$ at each point $x$ is obtained as $\hat{g}_{A}(x)= \frac{1}{N}\Sigma^N_{i=1} \frac{1}{s_i} \hat{g}_{a}(\frac{x}{s_i})$. \footnote{To perform the deconvolution, the choice of two tuning parameters is required: the order of the spline and the penalty parameter of the first-step maximum likelihood procedure. The latter is optimally calibrated to obtain a variance matching the bias-corrected estimate in Table IV of kline2022systemic.}

Estimation and Policy Implications

We investigate whether the differences between IW and Efron matter in terms of estimation and policy implications. For instance, suppose a counselor's goal is to guide applicants on whether to avoid sending their applications to firms that are identifies as highly discriminatory or discriminatory-- those exceeding a specified contact gap threshold, such as 0.05 or 0 respectively. We thus calculate $\widehat{\text{Prob}}\left(\hat{Y}^k_{i,T}>0.05\right)$ and $\widehat{\text{Prob}}\left(\hat{Y}^k_{i,T}>0\right)$ at $T=5$, for $k \in \{\text{Efron, IW-MR}\}$.\footnote{These probabilities are calculated producing 1-period ahead forecasts at $T=5$ of contact gaps based on in-sample data from all five waves.} We find that the probability of classifying a firm as highly discriminatory (discriminatory) is 4.29% (64.29%) for IW compared to 1.43% (60%) for Efron, suggesting a higher degree of discrimination and thus different policy implications when using IW instead of Efron.

Forecasting

We compare the forecasting performance of IW relative to that of Efron and JS. For each wave $T= 3, 4$, we produce one-step-ahead forecasts of (unstudentized) contact gaps for each firm by the following methods: TS, which uses the time-series mean of contact gaps at time $T$; Pool, which uses the pooled mean at time $T$ and IW-MR. For Efron, we obtain forecasts as posterior mean estimates of the RE.\footnote{We use the code provided by kline2022systemic that produces Figure A13 in their paper, where they assess the out-of-sample forecast accuracy of the posterior means. We adapt the code to use data from waves $1, .., T$ to produce the forecast at $T= 3, 4$.} We then compare the out-of-sample forecasts from each method $k$, $\{\hat{Y}^k_{i,T}\}$ to the actual realizations $\{{Y}_{i,T+1}\}$, for the waves $4,5$. For each forecasting method $k$ and each firm $i$, the MSFE over the out-of-sample period is

equation*[equation* omitted — 86 chars of source]

Figure (ref) reports the difference $\Delta MSFE$ between the MSFE of forecasts for IW-MR and those for Efron for each firm. The horizontal axis shows the value of the gender contact gap at $T=4$. The figure reveals that Efron tends to make larger errors (the dots fall below zero) for firms that at the time of forecasting fell in the right tail or near the center of the distribution, illustrating a possible tyranny of the majority phenomenon (which could also be due to misspecification of the normality assumption and/or to the effect of the data-dependent choice of regularization parameter used by kline2022systemic when implementing Efron).

figure[figure omitted — 225 chars of source]

We also evaluate the aggregate performance of the different methods by reporting the average MSFE across $i$ in Table (ref), which shows that IW-MR is the best method.

table[table omitted — 381 chars of source]

Robustness of IW

To illustrate the robustness of IW vs. Efron, we conduct a subsampling exercise. We randomly draw $B=1000$ subsamples without replacement, each subsample $b$ consisting of $n_b=20$ firms. For each method $k$, where $k \in \{\text{Efron, IW-MR}\}$, and for each subsample $b$, we calculate the aggregate out-of-sample RMSFE: $$\text{RMSFE}_{b, k} = \sqrt{\frac{1}{n_b} \sum_{i=1}^{n_b} \left(Y_{i, T+1}-\hat{Y}^k_{i,T}\right)^2} $$ at $T=4$, with $\hat{Y}^\text{Efron}_{i,T}$ obtained using the full sample of firms. In Table (ref), we report the minimum, maximum, mean, median and 90th percentile of $\text{RMSFE}_{b, k}$ across the $B=1000$ samples for each method.

table[table omitted — 426 chars of source]

The results in Table (ref) demonstrate a sizeable reduction in the worst-case performance under IW compared to Efron. While the aggregate performance (i.e., the Mean or Median columns in the table) is slightly better for IW but comparable between the two methods, the difference in the 90% percentile and Max RMSFE is noticeable: the Max RMSFE is .09312 for Efron versus .08723 for IW, representing a 6.33% improvement. This indicates that Efron is more prone to poor performance depending on the composition of the subsample, whereas IW effectively reduces the risk of large RMSFE values. These findings highlight the robustness of our approach, as evidenced by its consistently strong performance across the state-space simulated through this subsampling exercise.

Forecasting Earnings

This section considers an out-of-sample exercise that applies IW to forecasting earnings residuals using the Panel Study of Income Dynamics (PSID).

Data

We consider earnings data from the PSID for 1968-1993.\footnote{We use data up to 1993 because from 1994 a major revision of the survey disrupted the continuity of PSID files, see kim2000notes. Moreover, after 1997 the PSID switched from an annual to a biannual data collection.} We follow the literature on income dynamics (e.g., meghir2004income) and select a sample of male workers, heads of household, aged between 24 and 55 (inclusive). We drop individuals identifying as Latino, with a spell of self-employment, with zero or top-coded wages and with missing records on race and education. We also require that the change in log earnings is not greater than $+5$ or less than $-3$. We consider earnings residuals obtained from a first stage panel data regression of log labor income of an individual $i$ at time $t$, $\tilde{Y}_{i,t}$, on education, a quadratic polynomial in age, race and year dummies. We denote by $Y_{i,t}$ the residuals from this regression.

The goal is to obtain individual one-year-ahead forecasts of earnings residuals $Y_{i,t}$.\footnote{Forecasting earnings residuals is of interest since they measure individual income risk. For instance, accurate forecasting of individual earnings residuals might be useful for prospective lenders when deciding on loan applications.}

Forecasting Performance

We compare the out-of-sample aggregate performance of IW-MR from Section (ref), versus using TS or Pool for all individuals.

We report results for the balanced samples consisting of $N=164$ ($N=790$) individuals with continuous earnings in all consecutive years for 1968-1993 (1968-1980). We further consider an unbalanced sample built using rolling windows of $T=3$ time periods of balanced samples of individuals (which delivers sample sizes ranging from $3960$ to $7912$). Forecasts are based on the model:

align[align omitted — 93 chars of source]

We use rolling windows of $T=2$ time periods and compare the out-of-sample forecasts from each method $k$, $\hat{Y}^k_{i,T}$, where $k \in \{\text{TS, Pool, IW-MR}\}$ to the actual realizations ${Y}_{i,T+1}$, for $t=1972,...,1992, i=1,...,N$.

For each forecasting method $k$ and each individual $i$, the MSFE over the out-of-sample period is

equation[equation omitted — 91 chars of source]

Table (ref) reports averages of $MSFE(k,i)$ across $i$ for each forecasting method $k$. The Table shows that, while TS clearly outperforms Pool in terms of average MSFE, IW further improves aggregate accuracy.

table[table omitted — 405 chars of source]

To gain some insight into which individuals benefit more from borrowing strength (i.e., are given higher weights to Pool by IW), in Figure (ref) we divide the N=164 individuals of the balanced sample into ten quantiles according to their lagged earnings (the vertical axis) for each year (the horizontal axis). Within each quantile we compute the forecasts given the higher weights by IW-MR: the size of the dots is proportional to the average weight attributed by IW-MR to Pool across individuals in that year and in that quantile.\footnote{We set the size option of the R package ggplot equal to the mean of weights attributed to Pool by the IW-MR rule for each quantile.} Figure (ref) shows that individuals near the median of the earnings residuals distribution are those who benefit from larger weights to Pool.

figure[figure omitted — 233 chars of source]

One possible interpretation of our findings is that in the PSID there is enough unobserved heterogeneity to make the TS forecast outperform Pool in the aggregate (as indicated by Table (ref)). However, an additional improvement in aggregate accuracy can be obtained by using IW, which tends to borrow strength for individuals near the median of the distribution. This finding confirms the usefulness of IW even in terms of aggregate performance.

Conclusion

Estimating random effects and forecasting with micropanels is challenging due to the short time dimension, and existing solutions have shortcomings. In this paper, we introduce a complementary approach that effectively addresses these limitations while imposing minimal assumptions. In practice, our method involves shrinking the time series mean towards the panel mean, utilizing individual-specific weights calculated solely from time series data. We propose three types of feasible weights: estimated oracle weights, Minimax Regret optimal weights, and inverse mean squared forecast error (MSFE) weights. Our findings indicate that Minimax Regret optimal weights offer superior performance, closely followed by (in-sample) inverse MSFE weights.

Our method is applicable to linear panel data models and value-added models, subject to the assumption that covariates have homogeneous coefficients. The extension to models with heterogeneous coefficients for covariates, as well as to alternative loss functions, is straightforward for the inverse-MSFE weights, which directly accommodate more general models and loss functions. Extending the derivation of Minimax Regret optimal weights to more general contexts is inherently more complex, and we reserve this exploration for future research. A further natural extension is to consider more general classes of shrinkage estimators, such as combinations of our individual-level shrinkage estimator and the James Stein estimator.

appendix\section{Proofs} \begin{proof}[\bf{Proof of Lemma (ref)}] The MSFEs for TS and Pool are immediate. For IW, first write \begin{align*} Y_{i,T+1} - \widehat{Y}^{IW}_{i,T} &= \left( Y_{i,T+1} - Y_{i,T} \right) {W}_{i,T-1} + (Y_{i,T+1} - \mu) (1- {W}_{i,T-1}). \end{align*} Then, under assumption (ref), we have that \begin{align*} & \mathrm{MSFE}(\mathrm{IW}, \theta_i) \\ &= \mathbb{E} \left[ \left( Y_{i,T+1} - \widehat{Y}^{IW}_{i,T} \right)^2 \right] \\ &= \mathbb{E} \left[ \left( Y_{i,T+1} - Y_{i,T} \right)^2 \right] \mathbb{E} \left[ {W}_{i,T-1}^2 \right] + \mathbb{E} \left[ (Y_{i,T+1} - \mu)^2 \left(1 - {W}_{i,T-1}\right)^2 \right] \\ &\;\;\; + 2 \mathbb{E} \left[ \left( Y_{i,T+1} - Y_{i,T} \right)(Y_{i,T+1}-\mu) {W}_{i,T-1} \left(1 - {W}_{i,T-1}\right) \right] \\ &= 2 \sigma_i^2 \mathbb{E} \left[ {W}_{i,T-1}^2 \right] + \mathbb{E} \left[ (A_i - \mu)^2 \left(1 - {W}_{i,T-1}\right)^2 \right] + \sigma_i^2 \mathbb{E} \left[ \left(1-{W}_{i,T-1} \right)^2 \right] \\ &\;\;\; + 2 \mathbb{E} \left[ \left( U_{i,T+1} \right)^2 {W}_{i,T-1} \left(1 - {W}_{i,T-1}\right) \right] \\ &= 2 \sigma_i^2 \mathbb{E} \left[ {W}_{i,T-1}^2 \right] + \mathbb{E} \left[ (A_i - \mu)^2 \left(1 - {W}_{i,T-1}\right)^2 \right] + \sigma_i^2 \mathbb{E} \left[ \left(1-{W}_{i,T-1} \right)^2 \right] \\ &\;\;\; + 2 \sigma_i^2 \mathbb{E} \left[{W}_{i,T-1} \left(1 - {W}_{i,T-1}\right) \right] \\ &= \sigma_i^2 + \sigma_i^2 \mathbb{E} \left[ {W}_{i,T-1}^2 \right] + \mathbb{E} \left[ (A_i - \mu)^2 \left(1 - {W}_{i,T-1}\right)^2 \right], \end{align*} which proves the lemma. \end{proof} Lemma (ref) is used to prove Theorem (ref). \begin{lemma} Let $\mathcal{M} = \{ \mathrm{TS}, \mathrm{Pool}, \mathrm{IW} \}$. Let Assumptions (ref) and (ref) hold. Then, \begin{align*} R(\mathrm{IW},\theta_i) \leq \sigma_i^2 \nu \end{align*} for each $\theta_i \in \Theta$, which is defined in (ref). Furthermore, the inequality above is strict if either $0 < {W}_{i,T-1} < 1$ or the inequality in (ref) is strict. \end{lemma} \begin{proof}[\bf{Proof of Lemma (ref)}] To bound $\mathrm{MSFE}(\mathrm{IW}, \theta_i)$, invoke Assumption (ref) to write \begin{align*} \mathbb{E} \left[ (A_i-\mu)^2 \left(1 - {W}_{i,T-1}\right)^2 \right] \leq \lambda_i^2 \mathbb{E} \left[ \left(1 - {W}_{i,T-1}\right)^2 \right]. \end{align*} This implies that \begin{align} \mathrm{MSFE}(\mathrm{IW}, \theta_i) &\leq \sigma_i^2 + \sigma_i^2 \mathbb{E} \left[ {W}_{i,T-1}^2 \right] + \lambda_i^2 \mathbb{E} \left[ \left(1 - {W}_{i,T-1}\right)^2 \right]. \end{align} Note that \begin{align*} R(\mathrm{IW},\theta_i) &= \max \left\{ 0, \mathrm{MSFE}(\mathrm{IW}, \theta_i) - \sigma_i^2 - \min \{ \sigma_i^2, \lambda_i^2 \} \right\}. \end{align*} If $\mathrm{MSFE}(\mathrm{IW}, \theta_i) < \sigma_i^2 + \min \{ \sigma_i^2, \lambda_i^2 \}$, then $R(\mathrm{IW},\theta_i) = 0$. In this case, there is nothing left to prove. Hence, it suffices to assume that $\mathrm{MSFE}(\mathrm{IW}, \theta_i) \geq \sigma_i^2 + \min \{ \sigma_i^2, \lambda_i^2 \}$. It follows from (ref) and Assumption (ref) that \begin{align} \begin{split} & \mathrm{MSFE}(\mathrm{IW}, \theta_i) - \sigma_i^2 - \min \{ \sigma_i^2, \lambda_i^2 \}\\ &\leq \sigma_i^2 \mathbb{E} \left[ {W}_{i,T-1}^2 \right] + \lambda_i^2 \mathbb{E} \left[ \left(1 - {W}_{i,T-1}\right)^2 \right] - \min \{ \sigma_i^2, \lambda_i^2 \} \\ &\leq \max \{ \sigma_i^2, \lambda_i^2 \} \left( \mathbb{E} \left[ {W}_{i,T-1}^2 \right] + \mathbb{E} \left[ \left(1 - {W}_{i,T-1}\right)^2 \right] \right) - \min \{ \sigma_i^2, \lambda_i^2 \} \\ &\leq \max \{ \sigma_i^2, \lambda_i^2 \} - \min \{ \sigma_i^2, \lambda_i^2 \} \\ &= (\lambda_i^2 - \sigma_i^2) \mathbb{I} ( \lambda_i^2 > \sigma_i^2 ) + (\sigma_i^2 - \lambda_i^2) \mathbb{I} ( \lambda_i^2 < \sigma_i^2 ) \\ &\leq \sigma_i^2 \nu, \end{split} \end{align} where $\mathbb{I}\{\cdot\}$ denotes the indicator variable and the third inequality uses the fact that ${W}_{i,T-1}^2 + \left(1 - {W}_{i,T-1}\right)^2 \leq 1$ if $0 \leq {W}_{i,T-1} \leq 1$. In conclusion, we have shown that $R(\mathrm{IW},\theta_i) \leq \sigma_i^2 \nu$ for each $\theta_i \in \Theta$. This proves the first conclusion of the lemma. The second conclusion follows from the facts that the inequality in (ref) will be strict if the inequality in (ref) is strict and that the third inequality in (ref) will be strict if $0 < {W}_{i,T-1} < 1$ with positive probability. \end{proof} \begin{proof}[\bf{Proof of Theorem (ref)}] We have \begin{align*} \min_{m \in \mathcal{M}} \mathrm{MSFE}(m, \theta_i) \leq \min_{m \in \{ \mathrm{TS}, \mathrm{Pool} \}} \mathrm{MSFE}(m, \theta_i) = \sigma_i^2 + \min \{ \sigma_i^2, \lambda_i^2 \}. \end{align*} Furthermore, the regrets for TS and Pool are \begin{align*} R(\mathrm{TS},\theta_i) &\geq \sigma_i^2 - \min \{ \sigma_i^2, \lambda_i^2 \}, \\ R(\mathrm{Pool},\theta_i) &\geq \lambda_i^2 - \min \{ \sigma_i^2, \lambda_i^2 \}. \end{align*} Note that \begin{align} \begin{split} \max_{\theta_i \in \Theta} R(\mathrm{TS},\theta_i) &\geq \max_{\theta_i \in \Theta} \left[ (\sigma_i^2 - \lambda_i^2) \mathbb{I} \{ \sigma_i^2 > \lambda_i^2 \} \right] = \sigma_i^2 \nu, \\ \max_{\theta_i \in \Theta} R(\mathrm{Pool},\theta_i) &\geq \max_{\theta_i \in \Theta} \left[ (\lambda_i^2 - \sigma_i^2) \mathbb{I} \{ \sigma_i^2 < \lambda_i^2 \} \right] = \sigma_i^2 \nu, \end{split} \end{align} where $\mathbb{I}\{\cdot\}$ denotes the indicator variable as before. The claim in Theorem (ref) then follows directly from Lemma (ref) and the inequalities in (ref). \end{proof} \begin{proof}[\bf{Proof of Theorem (ref)}] If $\lambda_i^2 = \sigma_i^2$, it follows from (ref) that \begin{align*} \mathrm{MSFE}(\mathrm{IW}, \theta_i) - 2 \sigma_i^2 & \leq 0, \end{align*} which proves the first conclusion of the theorem. As in Lemma (ref), the second conclusion follows from the facts that the inequality in (ref) will be strict if the inequality in (ref) is strict and that the third inequality in (ref) will be strict if $0 < {W}_{i,T-1} < 1$ with positive probability. \end{proof} \begin{proof}[\bf{Proof of Lemma (ref)}] First, consider the MSFE for the TS forecast. Write \begin{align*} \mathbb{E} \left[ \left( Y_{i,T+1} - Y_{i,T} \right)^2 | Y_{i,1},...,Y_{i,T-1} \right] &= \mathbb{E} \left[ \left( U_{i,T+1} - U_{i,T} \right)^2 | Y_{i,1},...,Y_{i,T-1} \right] \\ &= \mathbb{E} \left[ \left( U_{i,T+1} - U_{i,T} \right)^2 \right] = 2 \sigma_i^2. \end{align*} Now consider the MSFE for the Pool forecast. Note that \begin{align*} &\mathbb{E} \left[ (Y_{i,T+1}-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] \\ &= \mathbb{E} \left[ \left( A_i-\mu + U_{i,T+1} \right)^2 | Y_{i,1},...,Y_{i,T-1} \right] \\ &= \mathbb{E} \left[ (A_i-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] + \mathbb{E} \left[ U_{i,T+1}^2 | Y_{i,1},...,Y_{i,T-1} \right] + 2 \mathbb{E} \left[ A_i U_{i,T+1} | Y_{i,1},...,Y_{i,T-1} \right] \\ &= \mathbb{E} \left[ (A_i-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] + \mathbb{E} \left[ U_{i,T+1}^2 \right] + 2 \mathbb{E} \left[ A_i \mathbb{E} \left[ U_{i,T+1} | Y_{i,1},...,Y_{i,T-1}, A_i \right] \right] \\ &= \mathbb{E} \left[ (A_i-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] + \mathbb{E} \left[ U_{i,T+1}^2 \right] + 2 \mathbb{E} \left[ A_i \mathbb{E} \left[ U_{i,T+1} \right] \right] \\ &= \mathbb{E} \left[ (A_i-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] + \sigma_i^2. \end{align*} To obtain the MSFE for IW, first write \begin{align*} Y_{i,T+1} -\widehat{Y}^{IW}_{i,T} &= \left( Y_{i,T+1} - Y_{i,T} \right) {W}_{i,T-1} + (Y_{i,T+1} - \mu) (1- {W}_{i,T-1}). \end{align*} Then, we have that \begin{align*} & \mathrm{MSFE}(\mathrm{IW}, \theta_i | Y_{i,1},...,Y_{i,T-1}) \\ &= \mathbb{E} \left[ \left( Y_{i,T+1} -\widehat{Y}^{IW}_{i,T} \right)^2 \Big| Y_{i,1},...,Y_{i,T-1} \right] \\ &= \mathbb{E} \left[ \left( Y_{i,T+1} - Y_{i,T} \right)^2 | Y_{i,1},...,Y_{i,T-1} \right] {W}_{i,T-1}^2 + \mathbb{E} \left[ (Y_{i,T+1} - \mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] \left(1 - {W}_{i,T-1}\right)^2 \\ &\;\;\;\; + 2 \mathbb{E} \left[ \left( Y_{i,T+1} - Y_{i,T} \right) (Y_{i,T+1} - \mu) {W}_{i,T-1} \left(1 - {W}_{i,T-1}\right) \right] \\ &= 2 \sigma_i^2 {W}_{i,T-1}^2 + \left\{ \mathbb{E} \left[ (A_i-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] + \sigma_i^2 \right\} \left(1 - {W}_{i,T-1}\right)^2 \\ &\;\;\;\; + 2 \mathbb{E} \left[ \left( U_{i,T+1} - U_{i,T} \right) \left( A_i + U_{i,T+1} \right) | Y_{i,1},...,Y_{i,T-1} \right] {W}_{i,T-1} \left(1 - {W}_{i,T-1}\right) \\ &= 2 \sigma_i^2 {W}_{i,T-1}^2 + \left\{ \mathbb{E} \left[ (A_i-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] + \sigma_i^2 \right\} \left(1 - {W}_{i,T-1}\right)^2 + 2 \sigma_i^2 {W}_{i,T-1} \left(1 - {W}_{i,T-1}\right) \\ &= \sigma_i^2 (1 + {W}_{i,T-1}^2 ) + \mathbb{E} \left[ (A_i-\mu)^2 | Y_{i,1},...,Y_{i,T-1} \right] \left(1 - {W}_{i,T-1}\right)^2, \end{align*} which proves the lemma. \end{proof} \begin{proof}[\bf{Proof of Theorem (ref)}] To minimize maximum regret, we set \begin{align*} \tilde{W}_{i,T-1}^2 = \tilde{W}_{i,T-1}^2 + \tilde{\zeta}_{i,T-1}^2 \left(1 - \tilde{W}_{i,T-1}\right)^2 - \frac{\tilde{\zeta}_{i,T-1}^2}{\tilde{\zeta}_{i,T-1}^2 + 1}, \end{align*} equivalently, \begin{align*} \tilde{W}_{i,T-1} = 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T-1}^2 + 1}}. \end{align*} To verify that $\tilde{W}_i$ is the solution, consider the case that \begin{align*} {W}_{i,T-1} > 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T-1}^2 + 1}}. \end{align*} Then, \begin{align*} (1 - {W}_{i,T-1})^2 < \frac{1}{\tilde{\zeta}_{i,T-1}^2 + 1}, \end{align*} which in turns implies that \begin{align*} \tilde{\zeta}_{i,T-1}^2 (1 - {W}_{i,T-1})^2 < \frac{\tilde{\zeta}_{i,T-1}^2}{\tilde{\zeta}_{i,T-1}^2 + 1}. \end{align*} Thus, maximum regret is $\sigma_i^2 {W}_{i,T-1}^2$, which is larger than the solution. Now consider the other case that \begin{align*} {W}_{i,T-1} < 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T-1}^2 + 1}}. \end{align*} Now maximum regret is \begin{align*} {W}_{i,T-1}^2 + \tilde{\zeta}_{i,T-1}^2 \left(1 - {W}_{i,T-1}\right)^2 - \frac{\tilde{\zeta}_{i,T-1}^2}{\tilde{\zeta}_{i,T-1}^2 + 1}. \end{align*} It remains to show that if ${W}_{i,T-1} < 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T-1}^2 + 1}}$, \begin{align*} {W}_{i,T-1}^2 + \tilde{\zeta}_{i,T-1}^2 \left(1 - {W}_{i,T-1}\right)^2 - \frac{\tilde{\zeta}_{i,T-1}^2}{\tilde{\zeta}_{i,T-1}^2 + 1} - \left( 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T-1}^2 + 1}} \right)^2 > 0. \end{align*} The left-hand side of the inequality above is minimized when $ {W}_{i,T-1} = \tilde{\zeta}_{i,T-1}^2 / ( \tilde{\zeta}_{i,T-1}^2 + 1), $ that is, ${W}_{i,T-1} = 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T-1}^2 + 1}}$. This is also the unique minimizer and plugging this value into the left-hand side of the inequality above yields 0. Thus, the left-hand side of the inequality above must be strictly positive if ${W}_{i,T-1} < 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T-1}^2 + 1}}$. Hence, we have proved the desired result. \end{proof} \section{IW for Estimation of RE} Rather than focusing on the forecasting problem discussed in the body of the paper, in this Appendix we consider the problem of estimating the RE $A_i$. \subsection{The Model} The model is: \begin{align} Y_{i,t} = A_i + U_{i,t}, i=1,...,N; t=1,...,T, \end{align} where $A_i \sim (\mu, \lambda_i^2)$ and $U_{i,t} \sim (0, \sigma_i^2)$. $A_i, U_{i,1},\ldots,U_{i,T}$ are random variables, whereas $\mu$, $\lambda_i^2$ and $\sigma_i^2$ are parameters. In other words, we take the frequentist approach. We make the following assumption. \begin{assumption}[Independence] $A_i, U_{i,1},\ldots,U_{i,T}$ are mutually independent. \end{assumption} \subsection{Optimality of IW} In this section we show conditions under which IW is Minimax Regret optimal relative to the time series estimator or the common mean, in a simplified setting where the weights and the time series estimators are independent. Suppose $T=2$ and consider the following estimators: \begin{align} Time series (TS): \widehat{Y}^{TS}_{i,T} &= Y_{i,2}, \\ Common mean (Pool): \widehat{Y}^{Pool}_{i,T} &= 0, \\ Shrinkage (IW): \widehat{Y}^{IW}_{i,T} &= \widehat{Y}^{TS}_{i,T}W_{i,(T-1)} + \widehat{Y}^{Pool}_{i,T}(1-W_{i,(T-1)}) = Y_{i,2}W_{i,1} . \end{align} The next lemma derives the MSEs of TS, Pool and IW when the estimand is $A_i$. \begin{lemma} Consider the three estimators above. Then under Assumption (ref) we have \begin{align*} \mathrm{MSE}(\mathrm{TS}, \theta_i) &= \sigma_i^2, \\ \mathrm{MSE}(\mathrm{Pool}, \theta_i) &= \lambda_i^2, \\ \mathrm{MSE}(\mathrm{IW}, \theta_i) &= \sigma_i^2 \mathbb{E} \left[ {W}_{i,1}^2 \right] + \mathbb{E} \left[ (A_i)^2 \left(1 - {W}_{i,1}\right)^2 \right]. \end{align*} \end{lemma} Lemma (ref) suggests that the trade-off between TS and Pool depends on the “signal-to-noise" ratio $\lambda_i^2/\sigma_i^2$: Pool dominates when the ratio is less than 1 and TS dominates when it is greater than 1. Let $\mathcal{M}$ include $\mathrm{TS}$, $\mathrm{Pool}$, and $\mathrm{IW}$. We define regret as \begin{align} R(m,\theta_i) := \mathrm{MSE}(m, \theta_i) - \min_{h \in \mathcal{M}} \mathrm{MSE}(h, \theta_i). \end{align} The Minimax Regret (MMR) criterion selects the estimator $m$ that minimizes the maximum regret \[ \max_{\theta_i \in \Theta} R(m,\theta_i), \] where $\Theta$ is the parameter space. To derive analytical results for IW, we impose the following key regularity condition. \begin{assumption}[Individual Weight] The individual weight ${W}_{i,1}$ satisfies $0 \leq {W}_{i,1} \leq 1$ and \begin{align} \mathbb{E} \left[ (A_i)^2 \left(1 - {W}_{i,1}\right)^2 \right] \leq \mathbb{E} \left[ (A_i)^2 \right] \mathbb{E} \left[ \left(1 - {W}_{i,1}\right)^2 \right]. \end{align} \end{assumption} \subsection{Minimax Regret Optimality of IW} In this section, we show the conditions under which IW is optimal in terms of Minimax Regret. We restrict our attention to the parameter space, where the signal-to-noise ratio $\lambda_i^2/\sigma_i^2$ ranges from $1-\nu$ to $1+\nu$ for some $0 \leq \nu < 1$. \begin{align} \Theta = \Theta (\nu) := \{ (\sigma_i^2, \lambda_i^2) \in \mathbb{R}_{+}^2: 1- \nu \leq \lambda_i^2/\sigma_i^2 \leq 1+\nu \}. \end{align} The next theorem shows that IW (uniquely) minimizes maximum regret among TS, Pool, and IW under Assumptions (ref) and (ref). \begin{theorem} Let Assumptions (ref) and (ref) hold. Then, \begin{align*} \max_{\theta_i \in \Theta} R(\mathrm{IW},\theta_i) \leq \min \left\{ \max_{\theta_i \in \Theta} R(\mathrm{TS},\theta_i), \max_{\theta_i \in \Theta} R(\mathrm{Pool},\theta_i) \right\}, \end{align*} where $\Theta$ is defined in (ref). Furthermore, the inequality above is strict if either $0 < {W}_{i,1} < 1$ or the inequality in (ref) is strict.\end{theorem} \subsection{Proofs} \begin{proof}[\bf{Proof of Lemma (ref)}] The MSEs for TS and Pool are immediate. For IW, \begin{align*} & \mathrm{MSE}(\mathrm{IW}, \theta_i) \\ &= \mathbb{E} \left[ \left( A_{i} - \widehat{Y}^{IW}_{i,T} \right)^2 \right] \\ &= \mathbb{E} \left[ \left( A_{i} - Y_{i,2} \right)^2 \right] \mathbb{E} \left[ {W}_{i,1}^2 \right] + \mathbb{E} \left[ (A_{i})^2 \left(1 - {W}_{i,1}\right)^2 \right] \\ &\;\;\; + 2 \mathbb{E} \left[ \left( A_{i} - Y_{i,2} \right)(A_{i}) {W}_{i,1} \left(1 - {W}_{i,1}\right) \right] \\ &= \sigma_i^2 \mathbb{E} \left[ {W}_{i,1}^2 \right] + \mathbb{E} \left[ (A_i)^2 \left(1 - {W}_{i,1}\right)^2 \right], \end{align*} which proves the lemma. \end{proof} Lemma (ref) is used to prove Theorem (ref). \begin{lemma} Let $\mathcal{M} = \{ \mathrm{TS}, \mathrm{Pool}, \mathrm{IW} \}$. Let Assumptions (ref) and (ref) hold. Then, \begin{align*} R(\mathrm{IW},\theta_i) \leq \sigma_i^2 \nu \end{align*} for each $\theta_i \in \Theta$, which is defined in (ref). Furthermore, the inequality above is strict if either $0 < {W}_{i,T-1} < 1$ or the inequality in (ref) is strict. \end{lemma} \begin{proof}[\bf{Proof of Lemma (ref)}] To bound $\mathrm{MSE}(\mathrm{IW}, \theta_i)$, invoke Assumption (ref) to write \begin{align*} \mathbb{E} \left[ (A_i)^2 \left(1 - {W}_{i,1}\right)^2 \right] \leq \lambda_i^2 \mathbb{E} \left[ \left(1 - {W}_{i,1}\right)^2 \right]. \end{align*} This implies that \begin{align} \mathrm{MSE}(\mathrm{IW}, \theta_i) &\leq \sigma_i^2 \mathbb{E} \left[ {W}_{i,1}^2 \right] + \lambda_i^2 \mathbb{E} \left[ \left(1 - {W}_{i,1}\right)^2 \right]. \end{align} Note that \begin{align*} R(\mathrm{IW},\theta_i) &= \max \left\{ 0, \mathrm{MSE}(\mathrm{IW}, \theta_i) - \min \{ \sigma_i^2, \lambda_i^2 \} \right\}. \end{align*} If $\mathrm{MSE}(\mathrm{IW}, \theta_i) < \min \{ \sigma_i^2, \lambda_i^2 \}$, then $R(\mathrm{IW},\theta_i) = 0$. In this case, there is nothing left to prove. Hence, it suffices to assume that $\mathrm{MSE}(\mathrm{IW}, \theta_i) \geq \min \{ \sigma_i^2, \lambda_i^2 \}$. It follows from (ref) and Assumption (ref) that \begin{align} \begin{split} & \mathrm{MSE}(\mathrm{IW}, \theta_i) - \min \{ \sigma_i^2, \lambda_i^2 \}\\ &\leq \sigma_i^2 \mathbb{E} \left[ {W}_{i,1}^2 \right] + \lambda_i^2 \mathbb{E} \left[ \left(1 - {W}_{i,1}\right)^2 \right] - \min \{ \sigma_i^2, \lambda_i^2 \} \\ &\leq \max \{ \sigma_i^2, \lambda_i^2 \} \left( \mathbb{E} \left[ {W}_{i,1}^2 \right] + \mathbb{E} \left[ \left(1 - {W}_{i,1}\right)^2 \right] \right) - \min \{ \sigma_i^2, \lambda_i^2 \} \\ &\leq \max \{ \sigma_i^2, \lambda_i^2 \} - \min \{ \sigma_i^2, \lambda_i^2 \} \\ &= (\lambda_i^2 - \sigma_i^2) \mathbb{I} ( \lambda_i^2 > \sigma_i^2 ) + (\sigma_i^2 - \lambda_i^2) \mathbb{I} ( \lambda_i^2 < \sigma_i^2 ) \\ &\leq \sigma_i^2 \nu, \end{split} \end{align} where $\mathbb{I}\{\cdot\}$ denotes the indicator variable and the third inequality uses the fact that ${W}_{i,1}^2 + \left(1 - {W}_{i,1}\right)^2 \leq 1$ if $0 \leq {W}_{i,1} \leq 1$. In conclusion, we have shown that $R(\mathrm{IW},\theta_i) \leq \sigma_i^2 \nu$ for each $\theta_i \in \Theta$. This proves the first conclusion of the lemma. The second conclusion follows from the facts that the inequality in (ref) will be strict if the inequality in (ref) is strict and that the third inequality in (ref) will be strict if $0 < {W}_{i,T-1} < 1$. \end{proof} \section{Feasible Weights for the General Case} This appendix derives the feasible weights reported in Section (ref) in the general case that $\textrm{TS} = \bar{Y}_{i,T}=\Sigma^T_{t=1}Y_{i,t}/T$, $\textrm{Pool}=\mu$ and $\textrm{IW}=\bar{Y}_{i,T}W_{i,T}+\mu(1-W_{i,T})$. \subsection{IW-O} The oracle weights that minimize the individual MSFE of IW are \begin{align} W^{o}_{i,T} = \frac{\lambda_i^2}{\lambda_i^2 + \sigma_i^2/T}. \end{align} These weights follow for example from equation (9) in Chapter 4 of timmermann2006forecast, using the fact that the joint distribution of $Y_{i,T+1}$ and $\widehat{Y}^{TS}_{i,T}$ is $$\begin{pmatrix} Y_{i,T+1}\\ \widehat{Y}^{TS}_{i,T} \end{pmatrix} \sim \left( \begin{pmatrix} \mu\\ \mu \end{pmatrix}, \begin{pmatrix} \lambda_i^2+\sigma_i^2 & \lambda_i^2\\ \lambda_i^2 & \lambda_i^2+\sigma_i^2/T \end{pmatrix} \right),$$ which gives the optimal weight on $\widehat{Y}^{TS}_{i,T}$ as the product between the inverse of the variance of the forecast and the covariance between the outcome and the forecast. The linear combination with $W^{o}_{i}$ as weight could also be obtained by applying the “best linear rule" in equation (9.4), page 129 of efron1973stein, to $\widehat{Y}^{TS}_{i,T}|A_i\sim(A_i,\sigma_i^2/T)$ and $A_i\sim(\mu,\lambda_i^2)$. Feasible oracle weights can for example be obtained as \begin{align} \frac{ \Sigma^T_{t=1} (Y_{i,t}-\mu)^2 /T -\Sigma^{T-1}_{t=1} (Y_{i,t} - Y_{i,t+1})^2 /2(T-1)}{\Sigma^T_{t=1} (Y_{i,t}-\mu)^2 /T -\Sigma^{T-1}_{t=1} (Y_{i,t} - Y_{i,t+1})^2 /2T }, \end{align} using the facts that: $\Sigma^T_{t=1} (Y_{i,t}-\mu)^2 /T$ is an unbiased estimator of $\lambda_i^2+\sigma_i^2$; the denominator of the oracle weights can be rewritten as $\lambda_i^2+\sigma_i^2-\frac{T-1}{T}\sigma_i^2$ and that $\widehat{\sigma}^2_i=\Sigma^{T-1}_{t=1} (Y_{i,t} - Y_{i,t+1})^2 /2(T-1)$ is an unbiased estimator of $\sigma_i^2$. Since both the numerator and the denominator in ((ref)) can be negative, we found that these feasible weights perform very poorly in simulations. We thus considered a number of alternatives and found that the best performance in simulations is obtained by taking the positive part of the numerator in ((ref)) and then again the positive part of the resulting weights, delivering the following feasible weights:\footnote{Alternatives such as just taking the positive part of the weights in ((ref)) or using the sample covariance between $Y_{i,t}$ and $Y_{i,t-1}$ as an estimator of $\lambda^2_i$ in the optimal weights numerator delivered very large errors in simulations.} \begin{align} W_{i,T}^{IW-O} = \left( \frac{ \left(\Sigma^T_{t=1} (Y_{i,t}-\mu)^2 /T - \Sigma^{T-1}_{t=1} (Y_{i,t} - Y_{i,t+1})^2 /2(T-1)\right)^+}{\Sigma^T_{t=1} (Y_{i,t}-\mu)^2 /T -\Sigma^{T-1}_{t=1} (Y_{i,t} - Y_{i,t+1})^2 /2T}\right)^+. \end{align} \subsection{IW-MR} To derive the IW-MR weights, we first derive the expression for the conditional MSFEs of TS, Pool and IW. \begin{lemma} Let Assumption (ref) hold. The MSFEs conditional on the information set at time $T$, $\mathcal{Y}_{N,T}$, are \begin{align*} \mathrm{MSFE}(\mathrm{TS}, \theta_i | \mathcal{Y}_{N,T}) &= \sigma_{i,T}^2 +\gamma_{i,T}^2, \\ \mathrm{MSFE}(\mathrm{Pool}, \theta_i | \mathcal{Y}_{N,T}) &= \sigma_{i,T}^2 + \kappa^2_{i,T} , \\ \mathrm{MSFE}(\mathrm{IW}, \theta_i | \mathcal{Y}_{N,T}) &= \sigma_{i,T}^2 +\gamma_{i,T}^2 {W}_{i,T}^2 +\kappa_{i,T}^2 \left(1 - {W}_{i,T}\right)^2 - 2\delta_{i,T}{W}_{i,T} \left(1 - {W}_{i,T}\right), \end{align*} where $\sigma_{i,T}^2 = \mathbb{E} \left[ U_{i,T+1} ^2 | \mathcal{Y}_{N,T} \right]$, $\gamma^2_{i,T}=\mathbb{E} \left[ \bar{U}_{i,T} ^2 | \mathcal{Y}_{N,T} \right]$, $\kappa^2_{i,T}=\mathbb{E} \left[ (A_i-\mu)^2 | \mathcal{Y}_{N,T} \right]$ and $\delta_{i,T}=\mathbb{E} \left[ (A_i-\mu) \bar{U}_{i,T} | \mathcal{Y}_{N,T} \right]$, with $\bar{U}_{i,T}=T^{-1} \Sigma^T_{t=1}U_{i,t}$. \end{lemma} \begin{proof}[\bf{Proof of Lemma (ref)}] The MSFE for the TS forecast is given by \begin{align*} \mathbb{E} \left[ \left( Y_{i,T+1} - \bar{Y}_{i,T} \right)^2 | \mathcal{Y}_{N,T} \right] &= \mathbb{E} \left[ \left( U_{i,T+1} - \bar{U}_{i,T} \right)^2 | \mathcal{Y}_{N,T} \right] \\ &= \mathbb{E} \left[ U_{i,T+1} ^2 | \mathcal{Y}_{N,T} \right] +\mathbb{E} \left[ \bar{U}_{i,T} ^2 | \mathcal{Y}_{N,T} \right] - 2 \mathbb{E} \left[ \bar{U}_{i,T} \mathbb{E} \left[ U_{i,T+1} | \mathcal{Y}_{N,T}, \bar{U}_{i,T} \right] | \mathcal{Y}_{N,T} \right]\\ &= \sigma_{i,T}^2 +\gamma^2_{i,T}, \end{align*} where the last equality follows from \begin{align} \mathbb{E} \left[ U_{i,T+1} | \mathcal{Y}_{N,T}, \bar{U}_{i,T} \right] = \mathbb{E} \left[ U_{i,T+1} | \mathcal{Y}_{N,T}, A_{i} \right] = 0. \end{align} Now consider the MSFE for the Pool forecast. Note that \begin{align*} \mathbb{E} \left[ (Y_{i,T+1}-\mu)^2 | \mathcal{Y}_{N,T} \right] &= \mathbb{E} \left[ \left( A_i -\mu+ U_{i,T+1} \right)^2 | \mathcal{Y}_{N,T} \right] \\ &= \mathbb{E} \left[ (A_i-\mu)^2 | \mathcal{Y}_{N,T} \right] + \mathbb{E} \left[ U_{i,T+1}^2 | \mathcal{Y}_{N,T} \right] + 2 \mathbb{E} \left[ (A_i-\mu) U_{i,T+1} | \mathcal{Y}_{N,T} \right] \\ &= \mathbb{E} \left[ (A_i-\mu)^2 | \mathcal{Y}_{N,T} \right] + \mathbb{E} \left[ U_{i,T+1}^2 | \mathcal{Y}_{N,T} \right] + 2 \mathbb{E} \left[ A_i \mathbb{E} \left[ U_{i,T+1} | \mathcal{Y}_{N,T}, A_i \right] | \mathcal{Y}_{N,T} \right] \\ &= \kappa^2_{i,T}+\sigma_{i,T}^2, \end{align*} where the last equality again follows from (ref). To obtain the MSFE for IW, first write \begin{align*} Y_{i,T+1} -\widehat{Y}^{IW}_{i,T} &= \left( Y_{i,T+1} - \bar{Y}_{i,T} \right) {W}_{i,T} + (Y_{i,T+1}-\mu) (1- {W}_{i,T}). \end{align*} Then, we have that \begin{align*} & \mathrm{MSFE}(\mathrm{IW}, \theta_i | \mathcal{Y}_{N,T}) \\ &= \mathbb{E} \left[ \left( Y_{i,T+1} -\widehat{Y}^{IW}_{i,T} \right)^2 \Big| \mathcal{Y}_{N,T} \right] \\ &= \mathbb{E} \left[ \left( Y_{i,T+1} - \bar{Y}_{i,T} \right)^2 | \mathcal{Y}_{N,T} \right] {W}_{i,T}^2 + \mathbb{E} \left[ (Y_{i,T+1}-\mu)^2 | \mathcal{Y}_{N,T} \right] \left(1 - {W}_{i,T}\right)^2 \\ &\;\;\;\; + 2 \mathbb{E} \left[ \left( Y_{i,T+1} - \bar{Y}_{i,T} \right)(Y_{i,T+1}-\mu) {W}_{i,T} \left(1 - {W}_{i,T}\right)| \mathcal{Y}_{N,T} \right] \\ &= \left\{\sigma_{i,T}^2 +\mathbb{E} \left[ \bar{U}_{i,T} ^2 | \mathcal{Y}_{N,T} \right]\right\} {W}_{i,T}^2 + \left\{ \mathbb{E} \left[ (A_i-\mu)^2 | \mathcal{Y}_{N,T} \right] + \sigma_{i,T}^2 \right\} \left(1 - {W}_{i,T}\right)^2 \\ &\;\;\;\; + 2 \mathbb{E} \left[ \left( U_{i,T+1} - \bar{U}_{i,T} \right) \left( A_i -\mu+ U_{i,T+1} \right) | \mathcal{Y}_{N,T} \right] {W}_{i,T} \left(1 - {W}_{i,T}\right) \\ &= \left[\sigma_{i,T}^2 +\gamma_{i,T} ^2 \right] {W}_{i,T}^2 + \left[\kappa^2_{i,T} + \sigma_{i,T}^2 \right] \left(1 - {W}_{i,T}\right)^2 \\ &\;\;\;\; + 2\left\{ \sigma_{i,T}^2 - \mathbb{E} \left[ (A_i-\mu) \bar{U}_{i,T} | \mathcal{Y}_{N,T} \right]\right\}{W}_{i,T} \left(1 - {W}_{i,T}\right) \\ &= \sigma_{i,T}^2 +\gamma_{i,T}^2 {W}_{i,T}^2 +\kappa_{i,T}^2 \left(1 - {W}_{i,T}\right)^2 - 2\delta_{i,T}{W}_{i,T} \left(1 - {W}_{i,T}\right), \end{align*} which proves the lemma. \end{proof} The optimal weights that minimize $\mathrm{MSFE}(\mathrm{IW}, \theta_i | \mathcal{Y}_{N,T})$ are \begin{align} W_{i,T}^* = \frac{\kappa_{i,T}^2+\delta_{i,T}}{\kappa_{i,T}^2 + \gamma_{i,T}^2+2\delta_{i,T}}. \end{align} We henceforth set $\delta_{i,T}\approx0$ (this can be seen as approximating $\mathbb{E} \left[ (A_i-\mu) \bar{U}_{i,T} | \mathcal{Y}_{N,T} \right]$ with the unconditional mean $\mathbb{E} \left[ ( A_i-\mu) \bar{U}_{i,T} \right]=0$). Define regret as the difference between the conditional MSFE for a generic weight $W_{i,T}$ and the conditional MSFE that corresponds to the optimal weights $W^*_{i,T}$ in ((ref)): \begin{align} R^*(W_{i,T}, \theta_{i,T} | \mathcal{Y}_{N,T}) &:= \mathrm{MSFE}(W_{i,T}, \theta_{i,T} | \mathcal{Y}_{N,T}) - \mathrm{MSFE}(W^*_{i,T},\theta_{i,T} | \mathcal{Y}_{N,T}) \\ &= \gamma_{i,T}^2 {W}_{i,T}^2 + \kappa_{i,T}^2 \left(1 - {W}_{i,T}\right)^2 - \frac{\gamma_{i,T}^2 \kappa_{i,T}^2}{\kappa_{i,T}^2 + \gamma_{i,T}^2} \nonumber \\ &= \gamma_{i,T}^2 \left[ {W}_{i,T}^2 + \zeta_{i,T}^2 \left(1 - {W}_{i,T}\right)^2 - \frac{\zeta_{i,T}^2}{\zeta_{i,T}^2 + 1} \right], \nonumber \end{align} where \begin{align} \zeta_{i,T}^2 := \frac{\kappa_{i,T}^2}{ \gamma_{i,T}^2}=\frac{\mathbb{E} \left[ (A_i-\mu)^2 | \mathcal{Y}_{N,T} \right]}{ \mathbb{E} \left[ \bar{U}_{i,T}^2 | \mathcal{Y}_{N,T} \right]}. \end{align} The following theorem obtains the optimal Minimax Regret weight under the assumption that we can put bounds on the random variable $\zeta_{i,T}^2$. \begin{theorem} Suppose that $\zeta_{i,T}^2$ in ((ref)) is such that $\zeta_{i,T}^2 \in [0, \tilde{\zeta}_{i,T}^2]$, where $\tilde{\zeta}_{i,T}^2$ is large enough that maximum regret can occur at $\zeta_{i,T}^2 = \tilde{\zeta}_{i,T}^2$. Consider maximum regret \begin{align*} \max_{\theta_i \in \Theta} R^*(W_{i,T}, \theta_i | \mathcal{Y}_{N,T}) &= {\gamma_{i,T}^2} \max \left[ {W}_{i,T}^2, \left\{ {W}_{i,T}^2 + \tilde{\zeta}_{i,T}^2 \left(1 - {W}_{i,T}\right)^2 - \frac{\tilde{\zeta}_{i,T}^2}{\tilde{\zeta}_{i,T}^2 + 1} \right\} \right], \end{align*} with $R^*(W_{i,T}, \theta_i | \mathcal{Y}_{N,T})$ defined as in ((ref)). Then, the weight that minimizes maximum regret is \begin{align} \tilde{W}_{i,T} = 1 - \frac{1}{\sqrt{\tilde{\zeta}_{i,T}^2 + 1}}. \end{align} \end{theorem} \begin{proof}[\bf{Proof of Theorem (ref)}] Same as proof of Theorem (ref) (with subscript $T$ instead of $T-1$). \end{proof} In applications, the value of the upper bound $\tilde{\zeta}_{i,T}^2$ is uncertain. We thus propose the following heuristic rule to obtain feasible Minimax Regret optimal weights for IW: \begin{align} \widehat{\tilde{\zeta}_{i,T}^2} := \frac{ \max \{ (Y_{i,1}-\mu)^2, ...,(Y_{i,T}-\mu)^2 \}}{ \Sigma^{T-1}_{t=1} (Y_{i,t} - Y_{i,t+1})^2 /2T(T-1)}. \end{align} Here, the denominator, $\Sigma^{T-1}_{t=1} (Y_{i,t} - Y_{i,t+1})^2 /2T(T-1)$, is an unbiased estimator of $\sigma_i^2/T$, which approximates $\gamma^2_{i,T}=\mathbb{E} \left[ \bar{U}_{i,T} ^2 | \mathcal{Y}_{N,T} \right]$ with the unconditional mean $\mathbb{E} \left[ \bar{U}_{i,T} ^2 \right]=\sigma^2_i/T$. The numerator, $\max \{ (Y_{i,1}-\mu)^2, ...,(Y_{i,T}-\mu)^2 \}$, is a proxy for the upper bound on $\kappa_{i,T}^2=\mathbb{E} \left[ (A_i-\mu)^2 | \mathcal{Y}_{N,T} \right]$.

\nocite{*}