Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,415 characters · 15 sections · 33 citation commands
Variance Estimation for Saturated Fixed-Effect Specifications
Applied econometric work routinely reports linear regressions with high-dimensional fixed effects: worker--firm models AKM, CHK2013, BLM2019, three-way trade gravity equations, triple-differences with cell structure, event studies with unit-and-time-interacted controls. As fixed effects (FE) saturate, two things happen at once: the proportional dimension of the FE space, $\rho_n = d_{K_n}/n$, grows large, and the residual treatment variance, $Q_K = \operatorname{\mathbb E}[\operatorname{Var}(X \mid \mathcal G_K)]$, shrinks. A natural conjecture, building on the weak-instrument literature SS, SY, Moreira2003, OleaPflueger2013, is that small $Q_K$ produces a weak-identification problem requiring its own critical values: identification analogous to a small concentration parameter.
The conjecture turns out to be false in the baseline model. Under strict exogeneity and conditional homoskedasticity, FE-residualized OLS is unbiased regardless of $Q_K$, because the within transformation is an exact orthogonal projection rather than an estimated first stage. The Cattaneo--Jansson--Newey-corrected $t$-statistic, defined with the residual-based variance estimator $\widehat\sigma^2 = \widehat u' \widehat u/(n - d_{K_n} - 1)$, has an exact $t_{n - d_{K_n} - 1}$ distribution; its size distortion relative to $N(0,1)$ is of order $1/n$ and does not depend on the residual treatment variance $\tau_n^2 = nQ_{K_n}$. There is no Stock--Yogo-style threshold in $\tau^2$ for this setup, in contrast to the instrumental-variables case. This negative result is a useful piece of econometric clarification, and it redirects attention to where the real distortion in saturated FE specifications does come from: variance estimation.
We characterize the asymptotic behavior of conventional variance estimators under the joint drift $(\rho_n, \tau_n^2) \to (\rho, \tau^2)$ with $\rho \in [0,1)$ and $\tau^2 \in (0, \infty]$. Three findings emerge.
First, the standard (no degrees-of-freedom correction) homoskedastic variance estimator $\widehat\sigma^2 = \widehat u' \widehat u / n$ produces $t$-tests with asymptotic size $2(1 - \Phi(z_{1-\alpha/2}\sqrt{1-\rho}))$, growing without bound in $\rho$. At $\rho = 0.5$ this is 17% for a nominal-5% test. The CJN correction $\widehat\sigma^2 = \widehat u' \widehat u/(n - d_{K_n} - 1)$ restores asymptotic exactness.
Second, the Eicker--White (HC0) heteroskedasticity-robust estimator has an asymptotic bias toward zero of order $\rho$: $\widehat\Omega_{\mathrm{HC0}}/(nQ_K) \to \omega^2(1-\rho) + \rho\,\operatorname{\mathbb E}[\sigma^2(X, \mathcal G_\infty)]$, where $\omega^2$ is the relevant limiting score variance. Under homoskedasticity this collapses to $\sigma^2$ (so HC0 is consistent in the absence of heteroskedasticity), but under genuine heteroskedasticity it produces $t$-tests with growing over-rejection.
Third, and less appreciated, the conventional HC3 estimator over-corrects in the opposite direction. Under homoskedasticity and balanced FE structure, $\widehat\Omega_{\mathrm{HC3}}/(nQ_K) \to \sigma^2 / (1-\rho)$, inflating the variance by factor $1/(1-\rho)$. This produces tests with over-coverage and corresponding under-rejection, with empirical size dropping below 1% at $\rho = 0.5$ for a nominal-5% test. HC3 is the standard recommendation when HC0 is suspected to be biased, but in saturated FE specifications it trades one error for the opposite error: HC0 under-covers the parameter, HC3 over-covers it.
The leave-one-out estimator of CJN --- equivalent to MacKinnon--White HC2 in the scalar-regressor case --- restores asymptotic exactness in both regimes. Under uniform leverage, \[ \widehat\Omega_{\mathrm{LO}} = \sum_i \widetilde X_{K_n, i}^2 \widehat u_i^2 / (1 - H_{ii}) \to \omega^2, \] the bias-corrected score variance. This recommendation is not new in principle: CJN originally proposed it for the many-covariates regime, and recent work by KSS uses related leverage-control techniques for variance components in bipartite FE models. Our contribution is to characterize the full HC0--HC1--HC2--HC3 hierarchy under the $(\rho, \tau^2)$ drift, document the HC3 over-correction theoretically and via simulation, and provide a practical recommendation grounded in the asymptotic theory.
We also note that an Anderson--Rubin-style robust test for the FE setting, $\mathrm{AR}_{K_n}(\beta_0) = (X' M_{K_n}(Y - X\beta_0))^2/\widehat\Omega_{\mathrm{LO}}(\beta_0)$, has uniform validity over $\tau^2 \in (0, \infty]$ under the maintained assumptions. Because the underlying score is unbiased under strict exogeneity, FE-AR coincides with the LO-based Wald confidence interval to high accuracy in our DGPs and offers no separate practical recommendation. We defer its statement and proof to Appendix (ref); the result is included because a companion paper on classical measurement error in the treatment uses the same construction in a setting where the score is biased and a non-trivial Stock--Yogo critical value emerges.
\paragraph{Related literature.} The conditional-expectation projection view of fixed effects is classical Davis, Wool. The AKM framework AKM has been extended to wage-inequality decomposition CHK2013 and to a distributional matched employer--employee framework BLM2019. CJN and JV study inference in the many-covariates regime ($\rho > 0$, $\tau^2 = \infty$ in our parametrization); KSS treat the AKM-bias-correction problem with leave-one-out techniques. The IV-analogue weak-identification literature SS, SY, Moreira2003, OleaPflueger2013, ASS, LMMP motivates the negative result on Stock--Yogo thresholds for FE saturation. GSU document the variance-weighted nature of saturated FE estimands; dCdH emphasize negative weights under heterogeneous effects, with parallel critiques in GoodmanBacon2021 for staggered-adoption DiD and an alternative aggregation approach in CallawaySantAnna2021. MacKinnonWhite1985, LongErvin2000, Imbens2016, and MacKinnon2023 survey the HC variance estimator family in standard (non-saturated) regression.
Section (ref) develops the setup, the population diagnostic $Q_K$, its sample analogue, and the joint CLT under the $(\rho, \tau^2)$ drift. Section (ref) characterizes the asymptotic behavior of conventional $t$-tests under both homoskedasticity (Section (ref), including the no-$\tau^2$-distortion result) and heteroskedasticity (Section (ref), the main contribution), with explicit limits for HC0, HC1, HC2 (LO), and HC3 expressed in terms of an effective design variance $\omega^2_{\mathrm{eff}}$. Section (ref) reports the simulation evidence. Section (ref) provides an empirical application: testing whether the Piotroski F-Score predicts forward equity returns in CEE markets, in a Visegrad firm-year panel where the theory's predictions about HC0/HC1/HC2/HC3 ratios can be verified against real data. Section (ref) concludes with a discussion of the sample concentration statistic $\widehat\tau_n^2$ as a descriptive diagnostic and the practical recommendation for applied work. Appendix (ref) records the Anderson--Rubin-style robust test.
Let $\{(Y_i, X_i)\}_{i=1}^n$ be a sample from a population on $(\Omega, \mathcal F, \mathbb P)$, with $Y_i \in \mathbb R$ and $X_i \in \mathbb R$ a scalar treatment. The structural model is
where $\mathcal G_\infty = \sigma(\bigcup_{K \geq 1} \mathcal G_K)$ is the limit of an increasing sequence $\mathcal G_1 \subseteq \mathcal G_2 \subseteq \cdots \subseteq \mathcal F$ of sub-$\sigma$-algebras representing successive FE specifications.
Let $D_K$ be the $n \times d_K$ matrix of FE dummies, $P_K = D_K (D_K' D_K)^{-1} D_K'$ the orthogonal projection, and $M_K = I_n - P_K$ the annihilator. Write $\widetilde X_K = M_K X$. The FE-residualized least-squares estimator is
The feasible residual from the full regression of $Y$ on $X$ and the FE dummies, computed by the Frisch--Waugh--Lovell theorem, is
where $H_X = \widetilde X_K (\widetilde X_K' \widetilde X_K)^{-1} \widetilde X_K'$ is the rank-one projection onto the within-transformed regressor. The residuals $\{\widehat u_i\}$ satisfy $\operatorname{\mathbb E}[\widehat u_i^2 \mid X, D] = \sigma_i^2 \cdot (1 - H_{ii})$ under independent errors, where $H_{ii} = (P_K)_{ii} + \widetilde X_{K, i}^2/(X' M_K X)$ is the leverage of observation $i$ in the full regression.
The population residual treatment variance is $Q_K = \operatorname{\mathbb E}[\operatorname{Var}(X \mid \mathcal G_K)]$, and the conditional error variance at point $i$ is $\sigma^2(X_i, \mathcal G_{K, i}) = \operatorname{\mathbb E}[u_i^2 \mid X_i, \mathcal G_{K, i}]$, with marginal mean $\sigma^2 = \operatorname{\mathbb E}[u^2]$. The asymptotic regime is governed by
The sample analogue is $\widehat Q_{K_n} = n^{-1} X' M_{K_n} X$ and the sample concentration statistic is $\widehat\tau_n^2 = n \widehat Q_{K_n} = X' M_{K_n} X$. With $\widehat\sigma_X^2 = n^{-1} X' M_0 X$ and $R_{K_n}^2$ the $R^2$ from regressing $X$ on the fixed effects, the equivalent representation is $\widehat\tau_n^2 = n \widehat\sigma_X^2 (1 - R_{K_n}^2)$. The within-cell sum of squares of the treatment, scaled by $n^{-1}$, plays the descriptive role of an identification diagnostic; we return to this in Section (ref).
We collect the standardized score, Hessian, and residual sum of squares and derive their joint limit.
Let $\widehat\sigma^2_{\mathrm{naive}} = \widehat u' \widehat u / n$ and $\widehat\sigma^2_{\mathrm{CJN}} = \widehat u' \widehat u / (n - d_{K_n} - 1)$, with corresponding $t$-statistics $T_n^{(j)} = (\widehat\beta_{K_n} - \beta_0)/(\widehat\sigma^{(j)}/\sqrt{X' M_{K_n} X})$.
Under heteroskedastic errors the limiting score variance $\omega^2$ may differ from $\sigma^2$. The leading estimators of the variance of $X' M_{K_n} u$ are:
where $H_{ii} = (P_{K_n})_{ii} + \widetilde X_{K_n, i}^2/(X' M_{K_n} X)$ is the full-regression leverage of observation $i$ defined in Section (ref), equation (ref). The $H_{ii}$ are the diagonal entries of the projection onto $\mathrm{span}(D_{K_n}, X)$, so $\sum_i H_{ii} = d_{K_n} + 1$ and the average full leverage is $(d_{K_n} + 1)/n$.
In the scalar-regressor case, the Cattaneo--Jansson--Newey leave-one-out variance estimator $\widehat\Omega_{\mathrm{LO}} = \sum_i \widetilde X_{K_n, i}^2 \widehat u_i \widehat u_i^{(-i)}$, where $\widehat u_i^{(-i)} = \widehat u_i/(1 - H_{ii})$ by Sherman--Morrison, reduces to $\widehat\Omega_{\mathrm{HC2}}$. We use the two labels interchangeably.
\paragraph{Two design quantities.} Two quantities of the design govern the asymptotic behavior of all four estimators under heteroskedasticity. The first is the target score-variance from Assumption (ref)(iv),
the $\widetilde X^2$-weighted average error variance. The second is the cross-leverage variance,
where $A := M_{K_n} - H_X$ is the projection associated with the residuals. The quantity $\bar\sigma_i^2$ is the average error variance over off-diagonal observations $j \neq i$, weighted by the squared cross-leverage $A_{ij}^2$. The two-quantity structure $(\omega^2, \mu)$ — rather than just $\omega^2$ alone — is what governs the HC limits below.
The effective design variance is the convex combination
with the proportional-FE weight $\rho$ controlling how much of the cross-leverage variance enters. Under conditional homoskedasticity, $\sigma_i^2 \equiv \sigma^2$ forces $\omega^2 = \mu = \sigma^2$ and $\omega^2_{\mathrm{eff}} = \sigma^2$. Under heteroskedasticity, $\omega^2_{\mathrm{eff}}$ equals the target $\omega^2$ exactly when $\mu = \omega^2$; we call this the cross-leverage balance condition. The condition is automatic under homoskedasticity and holds in a wider class of “design-balanced” heteroskedastic settings discussed in Remark (ref) below; under generic heteroskedasticity with non-uniform leverage, $\omega^2_{\mathrm{eff}} \neq \omega^2$ and an additional bias of order $\rho |\mu - \omega^2|$ enters every HC estimator. The simulation evidence in Section (ref) indicates that this bias is small in practice across the designs we test, but the theoretical statement below makes the dependence on $\mu$ explicit.
The proof is built from a concentration lemma and a conditional-expectation expansion, both of independent interest.
Appendix (ref) reports four Monte Carlo experiments validating the theoretical predictions. All four use balanced two-way fixed-effect panels with $X_{it} = \alpha_i + \gamma_t + \sigma_\eta\eta_{it}$ and $Y_{it} = \beta_0 X_{it} + a_i + b_t + u_{it}$, with $\beta_0 = 0$, $\sigma_\eta$ tuned to target $\tau^2 \in \{10, 100\}$, $n = 2000$, and $R = 5000$ replications per design cell. Q2 (naive vs.\ CJN $t$-test under homoskedasticity) confirms that the residual-DOF-corrected $t$-test has empirical size at the nominal $5\%$ level for $\rho \in \{0.05, 0.10, 0.25, 0.50\}$, while naive (uncorrected) tests over-reject at $\rho = 0.5$ with empirical size $9.5\%$. Q4 (within-$\widetilde X$ heteroskedasticity) and Q5 (one-way unbalanced FE producing non-uniform leverage) extend the variance-estimator predictions to settings the homoskedastic Q3 DGP does not probe. Across every cell of every experiment, empirical size is statistically indistinguishable from its value at $\tau^2 = 10$ vs $\tau^2 = 100$, confirming that the studentized test statistic is $\tau^2$-free: there is no Stock--Yogo-style threshold to enforce under strict exogeneity.
Table (ref) reports the headline experiment, Q3 (HC variants under heteroskedasticity), which validates the HC hierarchy of Theorem (ref) across the full saturation range.
The HC0 over-rejection at $\rho = 0.5$ is more than three times nominal; the HC3 under-rejection drops to a tenth of nominal. Both are practically meaningful in the applied range commonly seen in saturated specifications: a published TWFE design at $\rho \approx 0.25$ using HC0 has empirical size near $9\%$ where $5\%$ was claimed, so a reported 95% confidence interval has true coverage closer to 91%. Switching to LO restores nominal coverage at the cost of a correspondingly wider interval --- the wider interval is the honest one. HC1 is numerically indistinguishable from LO in this balanced design at every $\rho$ tested, consistent with the uniform-leverage prediction; the experiments Q4 and Q5 in Appendix (ref) break this equivalence under unbalanced designs.
We illustrate the variance hierarchy in a setting that motivates the small-$N$ panel regime this paper targets: testing whether the Piotroski F-Score Piotroski2000 predicts forward equity returns in Central and Eastern European (CEE) emerging markets, where listed-firm counts are modest and the FE saturation $\rho$ is consequently non-trivial in standard panel specifications.
The Piotroski F-Score is a 0--9 integer composite of nine binary accounting signals: four profitability indicators, three leverage/liquidity indicators, and two operating-efficiency indicators. The original Piotroski2000 study documented a long-high-short-low strategy on the U.S.\ Compustat universe, building on the expectations-errors-in-value/glamour-strategies tradition of LaPorta1996 and refined in PiotroskiSo2012; a complementary composite for growth firms is developed in Mohanram2005. Portability of the F-Score strategy to European markets has been investigated by Tikkanen2018, who find positive but heterogeneous returns to value-style F-Score sorts depending on the country mix and sample window; evidence specifically for emerging-market CEE settings remains sparse.
We work with a hand-collected Visegrad-country panel covering the Warsaw Stock Exchange (WSE, Poland), the Budapest Stock Exchange (BSE, Hungary), and the Prague Stock Exchange (PSE, Czech Republic) over fiscal years 2010--2024, including both active and delisted firms to address survivorship bias. After dropping firm-years missing either the F-Score or the one-year-ahead total return, the working sample has $n = 217$ firm-year observations across $N = 19$ unique firms and $T = 15$ years.
The point estimate of interest is the coefficient $\beta$ in the two-way fixed-effect regression
where $R_{i, t+1}$ is the realized one-year total return following the fiscal year-$t$ statements, $F_{i, t} \in \{0, 1, \ldots, 9\}$ is the Piotroski F-Score, and $\alpha_i, \gamma_t$ are firm and fiscal-year fixed effects. The within-FE residualized treatment $\widetilde X_{K_n}$ has $\widehat\tau_n^2 = X'M_K X \approx 559$ on the pooled sample, easily ruling out weak-identification concerns; the proportional FE dimension is $\rho_n = (N + T - 1)/n = 33/217 \approx 0.152$, moderate but non-trivial. This places the application in the regime that motivated the corrected Theorem (ref): $\rho$ is large enough that the HC0--HC3 spread should be visible at the second decimal of the standard error.
We report seven specifications. Specification (A) is the baseline of equation (ref) with arithmetic returns. The Hungarian observation 4IG-2017 (an enterprise-software roll-up that experienced a $\approx 1437\%$ return between fiscal year-end 2017 and year-end 2018) exerts substantial influence on (A) through the residual, but does not appear among the top-leverage observations because its F-Score (5) is essentially at the panel mean (5.81), producing a near-zero $\widetilde X_i$ after residualization. The largest-leverage observations in (A)--(C) are firms whose F-Score deviated sharply from its firm-year mean (CZG in Czech Republic, $\widetilde X = -2.72$ in 2020 and $+3.15$ in 2021; PLM.WA in Poland, $\widetilde X = -2.61$ in 2020 and $+2.27$ in 2021), at uniform $H_{ii} \approx 0.26$. (B) uses one-plus-log returns instead, and (C) winsorizes raw returns at the 1% and 99% quantiles, with (B) preferred. Specification (D) replaces the continuous F-Score with the indicator $\mathbf{1}\{F_{i,t} \geq 7\}$ (the original Piotroski long-side cutoff); (E) restricts to the Polish subsample (WSE only) at higher $\rho = 0.205$; (F) restricts to firms active throughout the sample (the survivorship-biased subsample) at $\rho = 0.154$; (G) replaces year FE with country-year FE, raising the saturation to $\rho = 0.290$ and the leverage spread to $h_{\max}/h_{\min} = 3.24$. The specifications (A)--(F) all maintain approximate leverage uniformity ($H_{\max} \leq 0.31$), while (G) generates four observations with $H_{ii} \geq 0.53$: PANNERGY and RICHTER in Hungary 2011, and CEZ and TABAK in Czech Republic 2012, the only firms observed in their respective country-year cells. The country-year FE absorbs the cell mean using only these two-observation cells, producing the leverage non-uniformity that drives the Spec G divergence between HC1 and LO discussed in Section (ref).
The point estimate $\widehat\beta$ is statistically indistinguishable from zero in every specification: the most favourable case for the F-Score is Spec E (Poland-only) with $\widehat\beta = +0.0135$ and SE$_{\mathrm{LO}} = 0.0220$, giving $|t| = 0.61$ and a one-sided $p$-value of approximately $0.27$. In every other specification $|t| < 0.35$. Independently of which variance estimator one selects, the Piotroski F-Score does not detectably predict forward returns in this Visegrad panel over 2010--2024. This is a useful finding by itself for the asset-pricing literature on CEE markets, but its principal role here is to set up the methodological comparison that follows: the inference-robustness exercise that one would naturally undertake to defend a marginal published finding does not, in this case, change the substantive conclusion --- but it does change other quantities by margins that are visible on first inspection and that match Theorem (ref) quantitatively.
Table (ref) compares the empirical ratios of the HC standard errors with the predictions of Theorem (ref) under the cross-leverage balance condition ($\omega^2_{\mathrm{eff}} = \omega^2$). The predicted ratios are $\mathrm{SE}_{\mathrm{HC0}}/\mathrm{SE}_{\mathrm{LO}} \to \sqrt{1-\rho}$, $\mathrm{SE}_{\mathrm{HC1}}/\mathrm{SE}_{\mathrm{LO}} \to 1$, and $\mathrm{SE}_{\mathrm{HC3}}/\mathrm{SE}_{\mathrm{LO}} \to 1/\sqrt{1-\rho}$.
The agreement between empirical ratios and the asymptotic predictions is striking at $\rho = 0.152$, where the gap is at most 0.011 in absolute value across specifications A--D, F. At higher saturation (Spec E, $\rho = 0.205$) the gaps remain below 0.01. At the most heavily saturated specification (G, $\rho = 0.290$), HC0/LO and HC3/LO each deviate by about 0.02 from theory, and the HC1/LO ratio rises to $1.031$ --- the only specification in which the uniform-leverage prediction visibly breaks. The mechanism is observable in the data: in (G), the country-year FE creates two-observation cells in Hungary 2011 (PANNERGY and RICHTER) and Czech Republic 2012 (CEZ and TABAK), where only two firms in the panel share a country-year. Each of the four observations attains $H_{ii} \approx 0.54$, well above the panel mean of $\rho_n = 0.29$, while the modal observation in (G) has $H_{ii}$ close to $0.29$. This is exactly the empirical signature of the leverage non-uniformity formalized in Remark (ref) and produced in the simulations of Q5 in Appendix (ref): a handful of high-leverage observations alongside a low-leverage modal mass, generating the slight HC1-over-LO bias that we observe.
The size of the LO-HC0 spread is economically meaningful. In specification G, switching from HC0 to LO widens the 95% confidence interval by 15%; switching from HC0 to HC3 widens it by 35%. In a setting where the published evidence on Piotroski returns in emerging markets is mixed, this is not a small adjustment. The substantive conclusion in our application is robust to the variance-estimator choice because $\widehat\beta$ is essentially zero; in applications where the published $\widehat\beta$ delivers $|t|$ between 2 and 3 under HC0, the LO correction can plausibly move the test below the conventional rejection threshold.
Three observations from the application generalize beyond it. First, the implied $\rho$ in CEE firm-year panels is non-trivial: in our pilot it reaches 0.29 once country-year FE are introduced, and would rise further with sector-by-year or firm-by-pre-post-event interactions. Standard inference under HC0 in such designs systematically under-states uncertainty. Second, the HC0--HC3 spread is consistent with the homoskedastic prediction even in a setting with substantial cross-sectional and time-series heterogeneity in return volatility; this is the empirical analog of the simulation finding (Q4) that the cross-leverage balance condition $\mu \approx \omega^2$ holds approximately in balanced symmetric designs. Third, when uniform leverage is approximately satisfied (specifications A--F here), HC1 is numerically indistinguishable from HC2/LO; once the design is genuinely unbalanced (Spec G), HC1 begins to deviate, exactly as predicted by Remark (ref).
The practical recommendation matches the one stated in the introduction: report $\widehat\beta$ with LO standard errors, accompanied by the diagnostic $\widehat\tau_n^2$ and the leverage spread $h_{\max}/h_{\min}$. The first communicates whether the within-FE identifying variation is adequate; the second communicates whether the LO advantage over HC1 is potentially material. In the Visegrad application both diagnostics are well-behaved, leaving the LO recommendation safely applicable.
Although the size of correctly studentized tests does not depend on $\tau^2$, the precision of the point estimate $\widehat\beta_{K_n}$ does: $\sqrt{nQ_{K_n}} (\widehat\beta_{K_n} - \beta_0) \Rightarrow N(0, \omega^2)$ by Corollary (ref), so the standard error scales as $\omega / \sqrt{\tau^2}$. A specification with $\widehat\tau_n^2 = 1$ yields a confidence interval of half-width $\approx 1.96 \omega$; a specification with $\widehat\tau_n^2 = 100$ yields half-width $\approx 0.20\omega$. We therefore recommend that authors using saturated FE specifications report $\widehat\tau_n^2 = n(1 - R_{K_n}^2)\widehat\sigma_X^2$ alongside the point estimate: it does not enter critical values --- there are none to enter, by Theorem (ref) --- but it communicates how much identifying variation in the treatment survives FE absorption, and so whether the estimate is sharp or fragile.
This paper traces the asymptotic behavior of conventional variance estimators in linear regression with saturated fixed effects, under a drift in which both the proportional fixed-effect dimension $\rho_n$ and the residual treatment variance $\tau_n^2$ vary with the sample size. The headline conclusions are: (i) the natural conjecture that small residual treatment variance produces weak-identification distortion fails in the baseline model, because FE-OLS is unbiased under strict exogeneity regardless of $\tau^2$; (ii) the conventional HC0 estimator is biased downward by a fixed factor $(1-\rho)$, producing over-rejection that grows with saturation; (iii) the conventional HC3 estimator over-corrects in the opposite direction by a factor $1/(1-\rho)$, producing under-rejection and a corresponding loss of power; (iv) HC2/leave-one-out removes the first-order leverage distortion that drives the HC0 and HC3 biases. HC2 is asymptotically exact under conditional homoskedasticity and under design-balanced heteroskedasticity --- the regime in which the cross-leverage variance $\mu$ coincides with the target score-variance $\omega^2$. Under general heteroskedasticity with non-uniform leverage or strong dependence between $\sigma_i^2$ and the within variation, HC2 retains an additional bias of order $\rho |\mu - \omega^2|$ characterized in Theorem (ref); the simulation evidence indicates this residual bias is small in balanced symmetric designs but can grow under unbalanced or asymmetric configurations.
The practical recommendation is straightforward: in saturated FE specifications, report $\widehat\beta_K$ with leave-one-out standard errors, accompanied by the descriptive diagnostic $\widehat\tau_n^2 = n(1 - R_{K_n}^2) \widehat\sigma_X^2$. Specifications with low $\widehat\tau_n^2$ have wide LO intervals, which correctly reflect the loss of identifying variation. Specifications using HC0 should be re-run with LO if heteroskedasticity is plausible; specifications using HC3 should be re-run with LO if $\rho > 0.1$, since HC3 over-correction will yield artificially conservative intervals. In designs with substantial leverage non-uniformity (extreme cell-size variation, “anchor” units with far more observations than the bulk), LO retains an edge over HC1 even when the latter is asymptotically valid under uniform leverage.
Three extensions are natural. The scalar-treatment restriction relaxes to vector treatment with a matrix-valued $Q_K$ and corresponding matrix-valued HC estimators; the smallest eigenvalue of $Q_K$ plays the role of the scalar $Q_K$ in our framework. The i.i.d.\ assumption can be replaced by clustered or network dependence, in which case the LO recipe must be replaced by an appropriate cluster-robust variant. Finally, the framework can be embedded in difference-in-differences with staggered adoption, where $\mathcal G_K$ encodes the treatment-cohort interaction structure and $\widehat\tau_n^2$ provides a finite-sample identification diagnostic complementary to the heterogeneous-effects critique of dCdH.
The author thanks Blessing Orjika and Dmitrii Verzun for their collaboration in collecting the Visegrad CEE firm-year panel used in Section (ref), originally assembled in the context of joint work on a capstone project. The author also thanks colleagues for detailed comments on multiple draft rounds.