Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,733 characters · 17 sections · 67 citation commands
Automatic Inference for Value-Added Regressions
{-5em}
\noindentKey words:\ \hangindent=1.8em empirical Bayes, shrinkage estimators, teacher value-added, \\errors in variables.
\noindentJEL classification codes: C12.
Empirical Bayes shrinkage estimators are widely applied in settings with a large number of latent individual effects observed via noisy measurements. {Economic applications include the study of teachers' value-added to student test scores chetty2014measuring,chetty2014measuring2, location effects chetty2018impacts2, hospital quality hull2018estimating, nursing home quality einav2025producing, firm-level discrimination kline2022systemic, among others.} Shrinkage estimates are typically not the final object, but instead serve as inputs to downstream analyses. A typical downstream use is to regress other variables of economic interest on the estimated individual effects.
It is common for researchers to treat the shrinkage estimates as if they were the true individual effects when performing conventional inference in the downstream regression. This widespread empirical approach implicitly assumes the property of automatic inference, i.e., no adjustment is required to achieve nominal coverage rates of confidence intervals. However, empirical Bayes shrinkage estimates generally differ from the latent true effects, introducing measurement error. Moreover, there is also the potential for a generated regressor problem pagan1984econometric, because the regressors are themselves generated in a first-stage estimation step. Thus, there may be threats to the validity of both the estimator (via measurement error bias), and inference (via incorrect standard errors).
In empirical applications, various linear shrinkage estimators have been used as regressors. A leading example is the literature studying how teachers' value-added for test scores impacts students' future outcomes. Data on students' test scores are aggregated into shrinkage estimates, which are constructed as a weighted average of each teacher's individual mean score and all teachers' pooled mean score. Most studies employ individualized shrinkage, where noisier estimates (e.g., from smaller classes) are shrunk more aggressively toward the pooled mean. {Empirical work using this approach includes jacob2007parents, jacob2008can, kane2008estimating, chandra2016health, jackson2018test, abdulkadirouglu2020parents, bau2020teacher, biasi2022flexible, warnick2024instructor, andrabi2025heterogeneity, and angelova2025algorithmic.} Individualized shrinkage differs from equal shrinkage approaches, such as chetty2014measuring,chetty2014measuring2, which apply a uniform shrinkage rule to all units. However, the theoretical impact of these different shrinkage schemes on downstream estimates and inference has not been studied to date.
This paper is the first to formally study the impact of different shrinkage schemes on the properties of downstream regressions, and to characterize when automatic inference does or does not obtain. We focus on linear shrinkage estimators, which are widely used in practice due to their simplicity and ease of implementation. We develop a general econometric framework for analyzing a broad class of individualized shrinkage estimators as regressors. These estimators are weighted averages whose weights reflect signal-to-noise ratios estimated from the data. They can be classified as individual-weight (individualized shrinkage) or common-weight (equal shrinkage). Our central finding shows that automatic inference hinges on the treatment of heteroskedastic measurement error. In particular, failing to account for heteroskedasticity in the shrinkage weights leads to invalid downstream inference, with asymptotically biased estimators and coverage rates from conventional confidence intervals below nominal coverage. By contrast, automatic inference obtains if the weights properly account for this heteroskedasticity. In this case, the conventional confidence intervals achieve nominal coverage without further correction, and the resulting downstream coefficient estimator is asymptotically equivalent to the infeasible OLS regression on the true latent effects. Thus, there is no efficiency loss from regressing on the estimated individual effects instead of the true individual effects.
These findings have important practical implications for current practice. Many empirical implementations model heteroskedastic measurement error solely by the number of measurements per individual, e.g., class sizes. Yet idiosyncratic noise variances may also differ across individuals. We show that accounting only for heterogeneity in the number of measurements can yield invalid inference if noise variances are correlated with the number of measurements. This is because the individualized adjustments systematically under- or over-correct depending on the degree of heterogeneity. This, in turn, introduces a non-classical measurement error, where the bias can lead to amplification rather than attenuation. Inference remains valid when the two sources of heterogeneity are independent. Properly accounting for heteroskedasticity in the shrinkage weights ensures robustness to such dependence and allows automatic inference for the downstream regression coefficient.
Accordingly, our analysis provides straightforward practical guidance for implementation. In the teacher value-added example, the weights may vary across teachers and are determined by the estimated signal-to-noise ratio. Components of the ratio are estimated by within- and across-teacher variances, which properly account for heteroskedasticity in the measurement error. In predicting long-run outcomes as a function of latent value-added, we simply regress the outcomes on these shrinkage estimates. From the reported results, conventional OLS standard errors and confidence intervals are valid for inference.
Our results do not impose distributional assumptions and have a clear intuition. Linear shrinkage estimators are the empirical Bayes posterior mean under normality of both the measurement error and the latent individual effect efron1973stein, morris1983parametric. However, while the functional form of shrinkage is derived under normality, we study regressions in a semiparametric context and none of our results rely on normality. The overall intuition is as follows. In the presence of noise, it is well known that regressions on unshrunk individual means suffer from a classical measurement error problem. Shrinkage estimators reduce the variability of the estimated individual effects and thereby offset the variance inflation that comes from noise. What is not clear is whether they offset it by just the right amount, so that downstream regression estimators are asymptotically unbiased and downstream inference is valid. We show that for some scenarios this is the case, while for others it is not.
To formalize our analysis, we first restate the established baseline result that the unshrunk individual mean suffers from classical errors-in-variables, resulting in attenuation bias and invalid inference when used as a downstream regressor. We then show our main results on individual-weight shrinkage. First, we outline the precise conditions under which inference fails, even when heterogeneity in the number of measurements is accounted for. Second, we establish how a properly constructed, fully individualized shrinkage estimator achieves asymptotic unbiasedness, inferential validity, and asymptotic efficiency in the downstream regression. Finally, we analyze the common-weight estimator as a point of comparison. Because of its close connection to conventional bias-correction methods, this estimator also addresses the inference problem. Nevertheless, we focus our attention on the individual-weight approach because it is more used in contemporary work and no theoretical results currently exist to justify its use in downstream regressions. While our focus has been on individual-weight shrinkage that assumes (as is standard) independence between individual effects and variance of measurement error, our results can be generalized to relax this assumption. Recent work has derived empirical Bayes estimators allowing such dependence chen2023mean, but has not used them in downstream regressions. In on-going work, we extend our analysis to this new class of estimators and show that they also yield automatic inference for downstream regressions.
We conduct Monte Carlo simulations to demonstrate the finite-sample performance of the different shrinkage estimators when used as regressors. The results confirm our asymptotic theory. The unshrunk individual means yield substantial attenuation bias and under-coverage of confidence intervals. Individual-weight shrinkage that accounts only for heterogeneity in the number of measurements can also lead to bias and under-coverage when noise variances are correlated with the number of measurements. In contrast, heteroskedastic individual-weight shrinkage that accounts for both sources of heterogeneity yields valid coverage, and achieves very close performance to regressing on the true latent individual effects. Common-weight shrinkage also performs well, but is outperformed by heteroskedastic individual-weight shrinkage in finite samples when the sample size is moderate.
We illustrate our results using two empirical applications. The first revisits the firm discrimination data of kline2022systemic, while the second is based on the school value-added data of andrabi2025heterogeneity. In the first application, we consider a regression of future discrimination levels on estimated discrimination at the firm level. We find that individual-weight heteroskedastic shrinkage yields confidence intervals more strongly supportive of a positive predictive effect. In the second application, we regress private school fees on school value-added estimates. We find that individual-weight heteroskedastic shrinkage produces slightly different confidence intervals from those reported in the original study, and it reaffirms the positive effect of school value-added on private school fees.
\paragraph{Related Literature} This paper relates to an older literature that studies the use of common-weight shrinkage estimators as downstream regressors. Regressing on common-weight shrinkage estimators can be traced back to whittemore1989errors, who demonstrates via simulations that common-weight shrinkage yields consistent downstream coefficients. That paper also points out informally its equivalence to bias-correction methods. Building on this insight, guo2012mean provide theoretical justifications regarding the quadratic risk reduction of such estimators for downstream coefficients. chetty2014measuring,chetty2014measuring2 construct common-weight shrinkage estimators based on the best linear predictor, emphasizing that the estimator is bias-free when used as the regressor. Their estimator can be viewed as equivalent to IV methods, and we unify this IV perspective along with the bias-correction perspective of common-weight shrinkage in (ref). deeb2021framework further exploits the equivalence to IV methods and develops inference results for chetty2014measuring,chetty2014measuring2's estimator, highlighting the need to adjust conventional OLS standard errors to account for errors in estimation of individual effects and nuisance parameters. While this line of work clarifies inference in the common-weight shrinkage setting, it essentially relies on the equivalence to IV and does not extend to individual-weight shrinkage. These results are not applicable to the analysis of individual-weight shrinkage, which we study.
We formalize and extend prior work on individual-weight shrinkage. Consistency of regression estimators based on individual-weight shrinkage has been briefly discussed in jacob2007parents, kane2008estimating, abdulkadirouglu2020parents, walters2024empirical, and andrabi2025heterogeneity. In particular, walters2024empirical highlights independence between individual effects and the variance of the noise as a necessary condition for using widespread individualized shrinkage strategies. Beyond such consistency results, there has been no work studying the consequences for downstream inference of using individual-weight shrinkage estimators as regressors. We show that valid inference depends critically on how heteroskedasticity is treated in the shrinkage weights. We establish asymptotic unbiasedness, asymptotic normality, and consistency of standard errors for the downstream regression estimator based on properly constructed individual-weight shrinkage.
This paper is also related to non-shrinkage options for various downstream models. Chang2024PostEB leverage information from the estimated prior to develop correction methods for nonlinear downstream models. For regressions, this literature strand is largely concerned with bias-correction. As noted from the equivalence, discussions from this category also apply to common-weight shrinkage. kline2020leave and bonhomme2024estimating emphasize correct estimation of moments for value-added that solves the errors-in-variables problem. chen2025regression show that standard bias-correction methods for the downstream regression remain consistent even if the individual effects and measurement error are correlated. We contribute to this literature by establishing conditions under which conventional shrinkage estimators yield valid inference for downstream regressions when individual effects and noise are independent. In ongoing work, we extend the analysis to this correlated case.
\paragraph{Outline}
This paper proceeds as follows. (ref) introduces the framework and its implications. (ref) discusses a broad class of shrinkage estimators, establishes their asymptotic normality and the consistency of standard errors, and discusses the validity of standard inference approaches for the downstream regression. It also provides some practical implementation guidelines. (ref) reports Monte Carlo simulations illustrating the finite-sample properties of the estimators. These results reinforce our main findings that properly constructed individualized shrinkage estimators yield valid inference, while improperly constructed ones do not. (ref) revisits the study of kline2022systemic on labor market discrimination. (ref) revisits the study of andrabi2025heterogeneity on school value-added. Finally, (ref) concludes. An appendix contains the additional lemmas, proofs, and extensions.
Each unit \(i\) is associated with a latent individual effect $\theta_i$, which we will refer to in what follows as value-added. Let $\bm\theta \coloneqq (\theta_i)_{i=1}^n$ denote the vector of these latent effects for the \(n\) sampled units. A large body of research studies the relationship between value-added $\theta_i$ and some outcome of economic interest $Y_i$. To fix ideas, consider the example of kane2008estimating. In that paper, $\theta_i$ is teacher $i$'s latent value-added for students' test scores, and $Y_i$ is the average long-term test score outcome of students taught by teacher $i$. The relationship is studied using the regression model
where the coefficient $\beta$ captures the downstream effect. It is assumed that the error term $u_i$ has mean zero and is uncorrelated with $\theta_i$, $\mathbb{E}\left[ u_i \theta_i\right] = 0$. Research questions about how teacher value-added relates to long term outcomes such as college attendance or earnings can be addressed by performing inference on $\beta$.
Inference on $\beta$ faces the challenge that $\bm \theta$ is unobserved, so the parameter of interest $\beta$ cannot be estimated by regressing $Y_i$ on $\theta_i$. Instead, for each unit $i$ we only have data $\left(Y_i, X_i\right) \coloneqq \left( Y_i, X_{i,1},...,X_{i,J_i}\right)$, $i=1,...,n$, where $Y_i$ is the outcome, $X_{i,1}, \ldots, X_{i, J_i}$ are repeated measurements of $\theta_i$, and $J_i$ is the number of observations available for estimating $\theta_i$. Specifically, $X_i \in \mathbb{R}^{J_i}$ are noisy repeated measurements for $\theta_i$ from the model
where $\epsilon_{i,j}$ is the noise term. In the example, $X_{i,j}$ is the test score of student $j$ taught by teacher $i$. The individual mean score for teacher $i$ is denoted by $\bar X_i$, representing the average test score within teacher $i$'s students. The pooled mean score is denoted by $\bar X$, representing the grand average across all teachers and students. In this paper, we primarily focus on the case where the noise is independent of the value-added, i.e., $\epsilon_{i,j} \perp\!\!\!\!\perp \theta_i$.
In this section, we analyze the inferential properties of different shrinkage estimators when used as downstream regressors. We follow the standard two-step workflow in empirical practice: Step 1: Estimate value-added $\theta_i$ with a shrinkage estimator $\hat \theta_i$; Step 2: Regress $Y_i$ on $\hat\theta_i$ and report conventional OLS estimates, standard errors and confidence intervals (i.e., treating the $\hat \theta_i$ as regular data rather than first-stage estimates). Our interest is in the validity of inference in Step 2 when different shrinkage estimators are used in Step 1.
The primary implication of our asymptotic framework is that the measurement error in $\bar{X}_i$ does not vanish relative to the sampling error for $\beta$, mimicking the finite-sample problem faced in practice, where both play a role. This persistence of measurement error is the central econometric challenge we address. As we will show in (ref), this asymptotic setup confirms that the naive unshrunk estimator (simply regressing $Y_i$ on $\bar{X}_i$) suffers from classical errors-in-variables bias and yields invalid inference. This provides a useful baseline and motivation for the other estimators.
We consider four primary classes of estimators for $\hat{\bm \theta}$, which we analyze in the order of our main results. For each class, we derive the asymptotic distribution of the OLS estimator of $\beta$ in the downstream regression. In (ref), we begin with the {unshrunk fixed-effect (FE)} estimator $\bar X_i$ as a baseline. We then move on to {individual-weight shrinkage}. We distinguish between two forms: {homoskedastic individual-weight shrinkage (HO)}, which models heterogeneity in measurement precision {solely by the number of measurements} ($J_i$), and {heteroskedastic individual-weight shrinkage (HE)}, which additionally models idiosyncratic noise variances. We study homoskedastic individual-weight shrinkage in (ref) and heteroskedastic individual-weight shrinkage in (ref). Finally, to connect with other approaches in the literature, we analyze {common-weight shrinkage (CW)}, which applies a single, uniform weight across units.
As a baseline, we first discuss using the fixed-effect estimator (FE) with no shrinkage. In this estimator, value-added for unit $i$ is simply estimated by the individual average: $\hat\theta_{i,\mathrm{FE}} = \bar X_i$, for $i=1,...,n$.
As we show formally below, using $\bar X_i$ as a regressor suffers from the classical errors-in-variables (EIV) problem, leading to attenuation bias. This bias shifts the location of standard OLS confidence intervals for $\beta$ away from the truth, rendering them invalid for inference. In the following, we first state and discuss assumptions that are needed for the asymptotic properties.
(ref).1 assumes independent numbers of measurements $J_i$. In (ref), the condition will be relaxed to allow dependence between $J_i$ and $\sigma_i^2$, accompanied by slightly stronger assumptions than (ref).7 and (ref).8. We keep the independence for ease of exposition. (ref).2 imposes unconditional exogeneity of the true $\theta_i$ and $u_i$, and moment conditions on $Y_i$.
(ref).3 places moment conditions on the noise $\epsilon_{i,j}$. The moment conditions are general and cover a wide range of distributions for $\epsilon_{i,j}/\sigma_i$, including normal, bounded, and even some asymmetric distributions. Thus, while normality of $\bar X_i$ is assumed to derive the functional form of parametric empirical Bayes estimators, the inference results we derive do not require any normality assumption. Independence $\epsilon_{i,j}\perp\!\!\!\!\perp \theta_i$ is standard in the empirical Bayes literature to justify shrinkage estimators without covariates. (ref).4 excludes further effects of measurement errors on $Y_i$ conditional on $\theta_i$. (ref).5 and (ref).6 are standard moment conditions for technical arguments.
Finally, (ref).7 and (ref).8 specify the asymptotic framework. Note that if we have the same number of measurements, $J_i = J$, then (ref).8 is implied by (ref).7. Thus (ref).7 is the essential condition. (ref) below holds without (ref).8. The asymptotic framework has three motivations. First, it aligns with the standard justification for normality of $\bar X_i$ via the central limit theorem with $J_i$ reasonably large walters2024empirical. Second, it reflects the data structure in common applications, such as the teacher value-added setting, where $J_i$ often has a similar magnitude to $\sqrt{n}$. In these settings, the number of teachers is in the thousands and the number of students per teacher is in the tens. For instance, in the study of North Carolina data in deeb2021framework, the total number of teachers $n = 5266$, and the total number of students is $388{,}191$, so we would expect $J_i$ to be on average about $74$, which is comparable to $\sqrt{n} \approx 73$. In another example from bau2020teacher, $\sqrt{n}\approx39$ and $J_i$ is on average about $15$. Third, it provides useful approximations to the finite-sample problem faced by researchers. The product $\sqrt{n}\mathbb{E} \left[J_i^{-1}\right]$ quantifies the relative scale of measurement error in the estimates of $\theta_i$ to sampling error in $\beta$. When $\kappa = 0$, $J_i$ grows asymptotically much larger than $\sqrt{n}$, implying that measurement error in $\theta_i$ is negligible for the purpose of inference in the regression. When $\kappa > 0$, measurement error in $\theta_i$ is of a similar order to the sampling error. In the previous two application examples, we have $\hat \kappa_1 \approx 73/74\approx1$, and $\hat \kappa_2 \approx 39/15\approx 2.6$,\footnote{These approximations are lower bounds, because Jensen's inequality produces a larger value if $J_i$ varies across $i$} as sample counterparts of $\kappa$. A finite positive $\kappa$ is thus appropriate for the context in which tens of students per teacher and thousands of teachers coexist.
Our asymptotic framework is related to the small-variance approximation in chesher1991effect and evdokimov2019errors,evdokimov2023simple. This literature does not put a structure on the form of the measurement error, but assumes its variance is proportional to $n^{-1/2}$. Here we know that the measurement error from estimation of $\theta_i$ has a variance proportional to $J_i^{-1}$, which leads to (ref).7. It is also standard in settings with both cross-sectional and individual-specific datasets, such as large $N$, $T$ panels pesaran2006estimation, factor augmented regressions gonccalves2014bootstrapping, and unstructured data battaglia2024inference.
Denote by $\hat\beta_{\text{FE}}$ the OLS estimator from the regression of $Y_i$ on $\hat\theta_{i,\mathrm{FE}}$ using observations $i= 1,...,n $. The following result implies that conventional OLS inference on $\beta$ using $\hat \beta_{\mathrm{FE}}$ is invalid.
In the special case where $\kappa=0$ (meaning measurement error is asymptotically negligible), the bias term disappears and the asymptotic distribution of $\sqrt{n} \left(\hat\beta_{\text{FE}} - \beta\right)$ is centered at zero. In our primary framework with $\kappa > 0$, $\hat\beta_{\mathrm{FE}}$ is consistent and asymptotically normal, with the efficient asymptotic variance. Thus, there is no generated regressor problem: OLS standard errors are consistent, following similarly to (ref) in (ref). However, its asymptotic distribution is centered away from zero due to attenuation bias. This has important practical consequences. Consider the conventional confidence interval for $\beta$, formed as $\hat\beta_{\mathrm{FE}} \pm 1.96 \times \mathrm{SE}(\hat\beta_{\mathrm{FE}})$ where $\mathrm{SE}(\hat\beta_{\mathrm{FE}})$ represents the usual OLS standard error. This confidence interval will be centered away from $\beta$, and consequently will not have valid coverage for $\beta$. To give a sense of the magnitude of under-coverage, the simulations in (ref) show that the coverage of a $95\%$ confidence interval is only about $70\%$.
In this subsection, we discuss the individual-weight shrinkage when the variance of $\bar X_i$ is assumed to only depend on the number of measurements $J_i$. That is, $\mathrm{Var}\left(\bar X_i \mid \theta_i\right) = \sigma^2/J_i$. Estimators constructed under this assumption are referred to in this paper as homoskedastic individual-weight shrinkage (HO), since the noise $\epsilon_{i,j}$ is homoskedastic across $i$ (note that we still allow for heteroskedasticity in the downstream regression, however). Here the primary source of heteroskedasticity is believed to arise from variation in the number of measurements across units $i$. The functional form of the shrinkage weights takes this heterogeneity into account.
Following the estimators proposed by jacob2007parents and kane2008estimating, the homoskedastic individual-weight shrinkage estimator $\bm{\theta}$ is constructed as
The estimator is the empirical Bayes posterior mean under normality of $\theta_i$ and $\epsilon_{i,j}$. The weight for $\bar X_i$ is $w_{i,\mathrm{HO}}$, obtained by replacing the oracle shrinkage weight \( \frac{\mathrm{Var}(\theta_i)}{\sigma^2/J_i + \mathrm{Var}(\theta_i)} \) with its empirical counterparts. These variances are estimated as
where $\bar X_{i,t}$, $\bar X_{i,t-1}$ are subset averages from splitting the measurements for each $i$ into two subsets. In a context with time periods, they can also be averages for two time periods. Note that the same $\hat\sigma^2$ is applied to units, with variation in $J_i$ reflecting differences in measurement precision.
We now derive the asymptotic distribution of the OLS estimator $\hat \beta_{\mathrm{HO}}$ of regressing $Y_i$ on $\hat\theta_{i,\mathrm{HO}}$ in two cases. In the first case, we assume the homoskedasticity in noise holds (i.e., $\sigma_i^2 = \sigma^2$ for all $i$). In the second case, we allow for heteroskedasticity in noise (i.e., $\sigma_i^2$ varies across $i$).
If homoskedasticity holds, this method is expected to correctly model the signal-to-noise ratio. Indeed, as we show below in (ref), regressing $Y_i$ on $\hat \theta_{i, \text{HO}}$ yields asymptotic unbiasedness and valid inference on $\beta$ under homoskedasticity.
If the true DGP is heteroskedastic, we demonstrate problems in inference when homoskedastic weights are applied. Here, denote the variance of $\bar X_i \mid \theta_i$ by $\sigma_i^2/J_i$. Regressing $Y_i$ on $\hat \theta_{i, \text{HO}}$ can lead to invalid inference on $\beta$ under noise heteroskedasticity. A leading case is when the number of measurements $J_i$ is correlated with the variance $\sigma_i^2$. In that case, it is natural to have $J_i$ endogenously increased for observations with large $\sigma_i^2$. But in this case the downstream regression estimator has an asymptotic bias which invalidates the standard approach to inference. The following result substantiates this claim, with details and proofs in (ref).
(ref) clarifies the conditions required for regressing $Y_i$ on $\hat\theta_{i,\mathrm{HO}}$ to yield valid inference. While inference is valid if the homoskedasticity assumption holds or if $J_i$ and $\sigma_i^2$ are independent, the estimator's performance is sensitive to violations of homoskedasticity. As $\gamma >0$ in the bias term, the asymptotic bias is amplification instead of classical EIV attenuation bias. Generally, the sign and magnitude of $\gamma$ depend on the dependence of $J_i$ and $\sigma_i^2$, and the bias can thus be positive or negative. It motivates the heteroskedastic estimator we consider next, which is designed to be robust to such dependence.
In this subsection, we discuss the {heteroskedastic individual-weight (HE) estimator} under noise heteroskedasticity. This estimator is designed to accommodate heterogeneity from both the number of measurements $J_i$ and the idiosyncratic noise variance $\sigma_i^2$. It improves upon $\hat\theta_{i,\mathrm{HO}}$.
The heteroskedastic individual-weight shrinkage estimator is defined as
Here, $\hat V$ is the estimator of $\mathrm{Var}\left(\theta_i\right)$, and $\hat \sigma_i^2$ is the estimator of $\mathrm{Var}\left(X_{i,j} \mid \theta_i\right)$. The variance estimators are constructed following kline2020leave and kline2022systemic:
Unlike $\hat\theta_{i,\mathrm{HO}}$, the weights $w_{i,\mathrm{HE}}$ allow $\mathrm{Var}(X_{i,j}\mid \theta_i)$ to differ across units, thereby ensuring robustness to heteroskedasticity.
Let $\hat\beta_{\mathrm{HE}}$ denote the downstream OLS estimator from regressing $Y_i$ on $\hat\theta_{i,\mathrm{HE}}$. We take an intermediate step in deriving the properties of $\hat\beta_\mathrm{HE}$. The estimator $\hat\theta_{i,\mathrm{HE}}$ is an empirical Bayes estimator, with the estimated prior variance $\hat{V}$ and prior mean $\bar X$ for a normal prior on $\theta_i$. We start by studying the property of OLS estimator $\hat\beta_{c,\mathrm{HE}}$ from regressing $Y_i$ on the semi-oracle estimator $\hat\theta_{i,c,\mathrm{HE}}$ (with prior mean still estimated from the data),
(ref) shows that the estimated variance $\hat V$ is a consistent estimator for the true variance $\mathrm{Var}\left(\theta_i\right)$. Thus, we build the properties of $\hat\beta_\mathrm{HE}$ based on those of $\hat\beta_{c,\mathrm{HE}}$. The following result shows that the OLS estimator $\hat \beta_{c,\mathrm{HE}}$ from regression $Y_i$ on $\hat\theta_{i,c,\mathrm{HE}}$ is asymptotically normal, correctly centered, with the efficient asymptotic variance from the infeasible regression of $Y_i$ on $\theta_i$.
(ref) shows that regressing $Y_i$ on $\hat\theta_{i,c,\mathrm{HE}}$ leads to an asymptotically unbiased estimator of $\beta$ in infeasible cases where $\mathrm{Var}\left(\theta_i\right)$ is known. The estimator performs asymptotically equivalently to that from regression of $Y_i$ on $\theta_i$. Next, we move on to the feasible estimator $\hat \theta_{i,\mathrm{HE}}$.
The following result builds on (ref) and establishes the asymptotic distribution if we use the feasible shrinkage estimator $\hat\theta_{i,\mathrm{HE}}$. Just like $\hat\beta_{c,\mathrm{HE}}$, the OLS estimator $\hat\beta_\mathrm{HE}$ from regressing $Y_i$ on $\hat\theta_{i,\mathrm{HE}}$ is also asymptotically normal, correctly centered, with the efficient asymptotic variance.
(ref) establishes that the OLS estimator based on the feasible regressor $\hat \theta_{i,\mathrm{HE}}$ is first-order asymptotically equivalent to the estimator based on the semi-oracle $\hat\theta_{i,c,\mathrm{HE}}$. Moreover, there is no efficiency loss relative to the infeasible regression of $Y_i$ on the latent $\theta_i$. The asymptotic variance is the efficiency bound for estimating $\beta$ in models where $\theta_i$ is observed chamberlain1987asymptotic. The asymptotic distribution doesn't depend on the value of $\kappa$. Thus, the OLS regression estimator $\hat \beta_{\mathrm{HE}}$ is robust to the amount of measurement error (under the conditions of (ref)).
We further establish consistency of standard errors from the regression of $Y_i$ on $\hat \theta_{i,\mathrm{HE}}$. Together with (ref), this ensures that regression of $Y_i$ on $\hat \theta_{i,\mathrm{HE}}$ delivers automatic inference in practice: one can proceed as if $\hat \theta_{i,\mathrm{HE}}$ are the true latent $\theta_i$ for the purpose of estimation and inference on $\beta$.
Building on (ref), additional regularity conditions are required for consistency of the variance estimator.
Here, (ref) strengthens the moment conditions to ensure that the law of large numbers applies to higher order terms in the variance estimator.
Taken together, (ref) and (ref) show that asymptotically valid confidence intervals for $\beta$ can be constructed as follows. If we regress $Y_i$ on $\hat\theta_{i,\mathrm{HE}}$, the reported standard error is $\sqrt{\hat{\Omega}/n}$. The asymptotically valid confidence intervals at level $1-\alpha$ are given by
where \( z_{1-\alpha/2} \) denotes the \( 1-\alpha/2 \) quantile of the standard normal distribution. (ref) and (ref) ensure that the confidence interval has asymptotically nominal coverage. Specifically:
In practice, efficient inference is easy to implement. One regresses $Y_i$ on $\hat{\theta}_{i,\mathrm{HE}}$ to obtain the OLS estimator $\hat{\beta}_{\mathrm{HE}}$ and its standard error. The resulting confidence interval for $\beta$ is asymptotically valid.
The HE estimator also serves another purpose. While the focus of this paper is on downstream inference, it is worth noting that the HE estimator has favorable properties for estimating the full vector of individual effects $\bm\theta$. As an estimator motivated by empirical Bayes principles, its construction improves estimation accuracy efron1973stein. By correctly modeling all sources of heterogeneity, $\hat{\bm\theta}_\mathrm{HE} \coloneqq \left(\hat\theta_{i,\mathrm{HE}}\right)_{i=1}^n$ is known to achieve a lower mean squared error (MSE) for $\bm\theta$ than the unshrunk FE, the HO, and the common-weight estimators, a standard result in that literature.
As a final point of comparison, we analyze the {common-weight (CW) shrinkage estimator}, which serves as a benchmark. This estimator, motivated by the James-Stein estimator, shrinks each unit mean $\bar X_i$ towards the grand mean $\bar X$ by a {common} factor $w$:
where $\bar X$ denotes the grand mean of the sample. The weight $w$ can be data-dependent as in the James-Stein estimator. For instance, common-weight shrinkage is applied in chetty2014measuring,chetty2014measuring2, where the weight $w$ is chosen from the best linear predictor of $\theta_i$ given $\bar X_i$. Inference of $\beta$ is then performed by regressing $Y_i$ on $\hat \theta_{i, \text{CW}}$.
We now give two examples of common weight $w$ for shrinkage.
\paragraph{(i) Bias Correction Shrinkage}
Regressing $Y_i$ on $\hat \theta_{i, \text{CW}}$ leads to an OLS estimator of $\beta$ given by
In effect, regressing $Y_i$ on $\hat \theta_{i,\text{CW}}$ adjusts $\hat\beta_{\text{FE}}$ of regressing $Y_i$ on $\bar X_i$ by a factor of $w^{-1}$. By a proper choice of shrinkage weight $w \approx \widehat \mathrm{Var}\left(\theta_i\right)/\widehat \mathrm{Var} \left(\bar X_i\right)$, the adjustment overlaps with the bias correction method for the EIV problem. Heuristically, in that case
The bias correction shrinkage method with the above $w$ thus has the equivalent property as classical bias correction in EIV problems for downstream regressions.
\paragraph{(ii) IV Shrinkage}
Another choice of $w$ would also result in a shrinkage option that, when used as the regressor, refines the IV estimator for $\beta$. To see this, if we split the measurements for each unit $i$ into two subsets, and treat the subset averages $\bar X_{i,1}$, $\bar X_{i,2}$ as the instrument and endogenous variable respectively, and then we construct an IV estimator:
The IV estimator adjusts the estimator from regressing $Y_i$ on $\bar X_{i,1}$ by a factor. With a proper modification of the factor, it can also adjust the estimator from regressing $Y_i$ on $\bar X_{i}$, which exploits more data.
Because of its connection to them, common-weight shrinkage can perform equivalently to the bias correction method and the IV estimator. It also requires as weak assumptions as theirs. Even though it ignores the heterogeneous signal-to-noise ratio and pools every individual at a uniform ratio, common-weight shrinkage provides valid inference on $\beta$. The following result formalizes this point for the bias correction shrinkage, with the variance estimator for $\theta_i$ the same as in $\hat\beta_\mathrm{HE}$. It is proved in (ref).
Therefore, the CW estimator also addresses the inferential problem from the FE baseline. Its primary distinction from the HE estimator is {how} it addresses the problem: the CW approach applies a uniform correction factor (equivalent to bias-correction or IV), while the HE estimator models unit-level heterogeneity. Our focus on individual-weight shrinkage is motivated by its prevalence in empirical applications. The common-weight approach characterizes influential work such as chetty2014measuring2. By contrast, a distinct and extensive literature has focused on individualized shrinkage jacob2007parents,kane2008estimating,andrabi2025heterogeneity. Our analysis thus provides the necessary theoretical foundation for this widely used empirical strategy.
In the first set of simulations, we focus on the comparison between regressing on the shrinkage estimator $\hat\theta_{i,\mathrm{HE}}$, the sample mean $\bar X_i$, i.e. $\hat\theta_{i,\mathrm{FE}}$, and the true latent $\theta_i$. In each of the $S = 3000$ simulations, we generate the data, run the regression of $Y_i$ on shrunk or unshrunk regressors, and then report the $95\%$ confidence intervals of $\beta$. Finally, for each value of $\beta$ in the grid, we compute the coverage rate across all simulations, i.e., the proportion of simulations in which the value of $\beta$ falls within the $95\%$ confidence interval.
The number of measurements is fixed at $J_i = J = 20$ and the sample size is $n = 1000$. Here the ratio $\hat \kappa \approx 1.58$. We set the true $\theta_i$ to be drawn from the standard normal distribution, and the variance $\sigma_i^2$ drawn from $\chi^2(1)$. In the regression, we set $\alpha = 0$, $\beta = 1$, and draw $u_i$ from a standard normal---thus homoskedastic---distribution.
In the first setting, we generate the measurement error $\epsilon_{i,j}$ from a normal distribution, with all conditions in (ref) satisfied. In the second setting, we set the measurement error $\epsilon_{i,j}$ still having the variance $\sigma_i^2$, but following a centered linear transform $\sigma_i\frac{\chi^2(2)-2}{2}$. Note that the ratio $\epsilon_{i,j}/\sigma_i$ satisfies the moment conditions in (ref). As $\epsilon_{i,j}$ follows a shifted Gamma distribution, which is non-normal and asymmetric, the purpose is to show that the proposed methods perform well under non-normal measurement error (provided the regularity conditions are satisfied).
In the second set of simulations, we compare bias and coverage rates across different methods, especially on the comparison between the individual-weight method HE and CW. We report the square root scaled MSE of $\hat\beta$, coverage rate, bias, and MSE of $\hat{\bm \theta}$. The design we consider lets the heteroskedastic $\sigma_i^2$ be from a uniform distribution on $\left\{1,10\right\}$, with other configurations the same as the previous case.
In the third set of simulations, we again compute the coverage rate for each method. We check the performance with random $J_i$ drawn from $\text{Poisson}(20)$, and with the sample size $n = 1000$. Here $\hat \kappa$ is approximately $1.66$, larger than $1.58$ in the first set of simulations due to convexity. With other parameters unchanged, now we can introduce correlation between $J_i$ and $\sigma_i^2$. In the first setting, we generate $J_i$ and $\sigma_i^2$ independently, ensuring that all conditions in (ref) are satisfied. In the second setting, we introduce a positive correlation between $J_i$ and $\sigma_i^2$, reflecting possible endogenous choice of more measurements for lower precision, keeping the conditions in (ref) satisfied.
The first set of coverage results is presented in (ref). Coverage for regressing $Y_i$ on $\hat \theta_{i, \mathrm{HE}}$ (the blue line) performs well in both normal and non-normal settings. The coverage rates are close to $95\%$ at the true $\beta$, and the curve is always very close to the infeasible case (the dashed red line) of regressing on the true $\theta_i$. In both cases, regressing on the sample mean $\bar X_i$ suffers from attenuation bias, reflected by a shift to the left in the coverage curve (the orange line), though the spread is approximately correct, which confirms the theory developed in (ref).
The second set of results is reported in (ref). The results indicate that regressing $Y_i$ on $\hat\theta_{i,\mathrm{HE}}$ has a smaller MSE of $\beta$, higher coverage rate, and smaller bias. Also, $\hat\theta_{i,\mathrm{HE}}$ has smaller MSE of $\bm\theta$ than $\hat\theta_{i,\mathrm{CW}}$ when the ratio $\sqrt{n}\mathbb{E}\left[J_i^{-1}\right]$ is reasonable. When $n$ and $J$ grow large, the performance of $\hat\theta_{i,\mathrm{CW}}$ and $\hat\theta_{i,\mathrm{HE}}$ becomes close, but $\hat\theta_{i,\mathrm{HE}}$ always dominates and is even more preferable when the sample size is relatively small. As discussed in (ref), we know that $\hat\beta_{\mathrm{CW}}$ is first-order asymptotically equivalent to $\hat\beta_\mathrm{HE}$. Combined with the simulation results, we can see that $\hat\beta_\mathrm{HE}$ can achieve better performance in finite samples.
The coverage results for random $J_i$ are presented in (ref). We mainly focus on the curves for $\hat\beta_\mathrm{HE}$ (the blue line) and $\hat\beta_\mathrm{HO}$ (the purple line). We can see that the $\hat\beta_\mathrm{HE}$ delivers valid coverage in both settings, close to the infeasible case (the dashed red line). Instead, $\hat\beta_\mathrm{HO}$ only delivers valid coverage in the first setting. When $J_i$ and $\sigma_i^2$ are correlated, the asymptotic distribution of $\hat\beta_{\text{HO}}$ is centered away from the true beta, reflected by a shift to the right in the coverage plot for $\hat\beta_{\text{HO}}$. Similar to before, regressing on the sample mean $\bar X_i$ results in a leftward biased coverage curve. As discussed in (ref), the asymptotic bias in $\hat\beta_\mathrm{HO}$ is distinct from classical EIV attenuation bias, as it shifts the estimate in the opposite direction and causes amplification.
Overall, the simulation results confirm the theoretical findings in (ref) and support (ref) regarding the nonnormality of $\epsilon_{i,j}$. Additionally, they validate the theoretical framework in (ref) concerning the dependence of $J_i$ and $\sigma_i^2$. The estimator $\hat\beta_\mathrm{HE}$ works well in both normal and non-normal settings, and is robust to the correlation between $J_i$ and $\sigma_i^2$. It outperforms $\hat\beta_\mathrm{FE}$, $\hat\beta_\mathrm{CW}$ and $\hat\beta_\mathrm{HO}$, and is close to the infeasible best case.
In this section, we illustrate the use of regression on heteroskedastic individual-weight shrinkage estimators in the context of kline2022systemic, which studies the extent to which large U.S. employers systemically discriminate job applicants based on race. Their study utilizes correspondence audits, where fictitious resumes with randomized racial identifiers are sent to employers to measure differences in callback rates. The racial contact gap, defined as the difference in callback probabilities between racial groups, serves as the primary latent variable, analogous to value-added in our setting.
We use the panel dataset in kline2022systemic from an experiment that sends fictitious applications to jobs posted by 108 of the largest U.S. employers. For each firm in each wave, about 25 entry-level vacancies were sampled and, for each vacancy, 8 job applications with randomly assigned characteristics were sent to the employer. Sampling was organized in 5 waves. Focusing on firms sampled in all waves yields a balanced panel of $n = 70$ firms over 5 waves. Applications were sent in pairs, one randomly assigned a distinctively White name and the other a distinctively Black name. The primary outcome is whether the employer attempted to contact the applicant within 30 days of applying. The racial contact gap is defined as the firm-level difference between the contact rate (the ratio of the number of contacts to the number of received applications) for White and that for Black applications. We follow the model similar to kline2022systemic, who assumes that the racial contact gap is given by
They justify the normality of $\bar \epsilon_{i}$ from the central limit theorem approximation with large numbers of measurements.
In our analysis, we estimate the predictive effect of callback probability in wave \( t \), for \( t = 1,2,3 \), on callback probability in wave \( t+2 \). Given the model setup, we expect the regression coefficient to be close to 1. For each wave \( t \), we first apply shrinkage estimation, and subsequently regress on these estimates. The shrinkage estimators constructed by kline2020leave are nonlinear shrinkage based on nonparametric empirical Bayes posterior means, while we focus on the performance of linear shrinkage estimators as regressors. We compare the performance of heteroskedastic individual-weight shrinkage estimator $\hat\beta_\mathrm{HE}$ with the fixed effect estimator $\hat\beta_\mathrm{FE}$ across the three waves. Since the discrimination gap is computed from pairs of job applications, we have \( J = 100 \), resulting in a small ratio of measurement error to sampling error, \( \sqrt{n}/J = 0.08 \).
The results presented in (ref) indicate that regression on the shrinkage estimates (\( \hat{\theta}_{i,\mathrm{HE}} \)) yields coefficients closer to 1. Despite a small \( \sqrt{n}/J \), $\hat\beta_\mathrm{FE}$ still exhibits attenuation bias. In terms of inference, $\hat\beta_\mathrm{HE}$ robustly rejects the null hypothesis of no predictive effect, whereas $\hat\beta_\mathrm{FE}$ fails to reject this null hypothesis for waves 1 and 2. These findings demonstrate that $\hat\beta_\mathrm{HE}$ improves the validity of inference for the regression coefficient.
In this section, we illustrate the use of regression on heteroskedastic individual-weight shrinkage estimators in the context of andrabi2025heterogeneity about the school value-added in Pakistan. Their study investigates the relation between private school fees and school value-added (SVA), and finds evidence that parents are able to identify and reward school quality.
The analysis utilizes the LEAPS project dataset, a rich longitudinal panel of student test scores from rural Punjab, Pakistan. This setting is characterized by the rapid emergence of a private school market, making the estimation of school quality (value-added) and its perceived return (school fees) a central question of interest. We use the school-year level sample from their analysis, which consists of $n=1158$ observations. We also have $\mathbb{E}_n \left[J_i^{-1}\right] = 0.120$. Therefore $\hat\kappa = \sqrt{n} \mathbb{E}_n \left[J_i^{-1}\right] \approx 4.08$.
We replicate and extend the primary downstream analysis in andrabi2025heterogeneity, which is a regression of private school fees on estimated SVA. The latent variable $\theta_i$ represents the true SVA of a school, and the downstream regression investigates whether schools with higher SVA command higher fees.
We compare two individualized shrinkage estimators. The first is HO, as used in andrabi2025heterogeneity. The second is HE, which is fully individualized by allowing for heterogeneity in both $J_i$ and the noise variance $\sigma_i^2$.
The results are presented in (ref) (full sample) and (ref) (a selected sample restricting $20 \leq J_i \leq 80$, as a robustness check under less heterogeneity). The tables compare the HE estimator (right column) and the HO estimator (left column) in both bivariate specifications and specifications including household-level controls (parental education and asset index).
Across all specifications, our results confirm the central economic finding of andrabi2025heterogeneity: SVA is a large, positive, and statistically significant predictor of private school fees. In the full-sample specification with controls ((ref), Row 2), the HE point estimate is 793.01 (significant at the 1% level), which is comparable to the HO estimate. This comparison provides a valuable robustness check. Since one cannot know a priori if the assumptions required for valid inference with $\hat\beta_{\mathrm{HO}}$ hold, the stability of the coefficients under the more general $\hat\beta_{\mathrm{HE}}$ estimator strengthens the credibility of the original economic conclusion.
This paper investigates the inferential properties of using individualized shrinkage estimators as covariates in a regression, a widely used but not fully understood empirical practice. We formalize the conditions under which this approach yields valid inference in the downstream regression model. Our central finding is that a correctly specified and fully individualized shrinkage estimator yields an asymptotically efficient estimator of the downstream regression coefficient. Crucially, we show that conventional OLS inference is asymptotically valid, justifying standard empirical practice. This result stands in contrast to the biased inference from regressing on the unshrunk fixed effect estimates (individual means). Similarly, the simpler individual-weight (HO) estimators which are widely used in practice can fail to deliver valid inference when noise variance is correlated with the number of measurements. Our simulations confirm the predictions of the theory, including the finding that regression on individualized shrinkage estimators that do not properly account for noise heteroskedasticity suffers from amplification bias. We apply our method to data on firm discrimination and school value-added and show that it improves the estimation and inference.
Our analysis provides a formal bridge between the empirical Bayes literature, where shrinkage estimators were developed to improve estimation accuracy of individual effects, and the common empirical practice of using these estimates for downstream inference. The key takeaway for practitioners is that while the plug-in approach can be valid and efficient, its robustness depends critically on the specification of the shrinkage estimator. For applied researchers, our results provide a clear theoretical foundation and a practical guide for obtaining valid inference when using shrinkage estimates as regressors in linear models. Extending this method to nonlinear settings is a natural yet nontrivial direction, which we are studying in ongoing work.