The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
155,729 characters
Fixed-Effect Saturation Is Not Weak Identification: Certifying Inference under Measurement Error
\begin{frontmatter}
\title{Fixed-Effect Saturation Is Not Weak Identification: Certifying Inference under Measurement Error}
\author{Stanis\l{}aw M. S. Halkiewicz~\orcidlink{0009-0000-7344-7522}}
\ead{[email removed]}
\affiliation{organization={Group of Machine Learning Research (GMUM), Jagiellonian University},
city={Krak\'ow},
country={Poland}}
\begin{abstract}
Fixed-effect saturation alone is not weak identification: in the baseline model, fixed-effect--residualized OLS is unbiased and conventional inference is asymptotically exact for every residual treatment variance $\tau^2=nQ_K>0$. Classical measurement error in the treatment restores it, and we derive Stock--Yogo-style critical values for $\tau^2$. Under the local drift $\sigma_\nu^2 = c^2/n$, attenuation produces a non-central limit whose non-centrality $\eta$ decreases in $\tau^2$ and depends on the fixed-effect dimension $\rho$ only through an overall $\sqrt{1-\rho}$ scaling, leaving the within reliability $\rho$-free. Inverting the leading quadratic size distortion gives a closed-form threshold; the breakdown reliability has a fixed-point form in the reported $t$-statistic alone. The diagnostic needs only a lower bound on reliability, where bias correction needs a point estimate. We separate a descriptive \emph{point pass} from a \emph{formal certificate}, evaluated at an upper confidence bound and carrying false-certification probability at most $\gamma$. A cluster-robust theory --- score CLT, Arellano-variance consistency under a checkable projection-compatibility condition --- yields $\eta_{CR}=\eta/\sqrt{\psi}$. Simulations confirm the threshold; in a saturated democracy--growth panel, aggregate V-Dem polyarchy is certified at $\gamma=0.05$ while its judicial-constraints sub-index is flagged under i.i.d.\ and clustered standard errors. The diagnostic covers classical error in a continuous regressor, \emph{not} binary-treatment misclassification.
\end{abstract}
\begin{keyword}
attenuation bias \sep local-to-zero asymptotics \sep panel data \sep weak identification \sep within reliability
\end{keyword}
\end{frontmatter}
\section{Introduction}\label{sec-intro}
The weak-instruments literature has formalized a familiar empirical concern: when the first-stage relationship between instruments and an endogenous regressor is weak, two-stage least squares is biased, and the bias does not vanish at the conventional $\sqrt n$ rate. \citet{SS} showed that the bias is governed by the concentration parameter $\mu^2$; \citet{SY} tabulated critical values for the first-stage $F$-statistic that bound this bias below pre-specified tolerances; \citet{LMMP} and \citet{ASS} refined the practical recommendations.
This paper poses an analogous question for linear regression with high-dimensional fixed effects. Empirical work routinely reports specifications with progressively richer FE structure: worker--firm models \citep{AKM}, gravity equations with three-way fixed effects, event studies with unit-and-time-interacted controls. As FE saturate, the residual variation of the treatment shrinks:
\[
Q_K = \operatorname{\mathbb E}[\operatorname{Var}(X \mid \mathcal G_K)] \to 0,
\]
where $\mathcal G_K$ is the $\sigma$-algebra generated by the $K$th FE specification. A natural conjecture is that small $Q_K$, like small $\mu^2$, produces a weak-identification problem requiring its own critical values.
The conjecture turns out to be false in the baseline model. Under strict exogeneity, FE-residualized OLS is \emph{unbiased} regardless of $Q_K$. With conditionally Gaussian homoskedastic errors, the Cattaneo--Jansson--Newey-corrected $t$-statistic (with residual variance formed after projecting out the treatment) has an exact $t_{n - d_K - 1}$ distribution; its size distortion relative to $N(0,1)$ is $O(1/n)$ and does not depend on $\tau^2 = nQ_K$. The IV analogy breaks at its foundation, because TSLS bias arises from the first-stage residual appearing in both the numerator and denominator of the second-stage estimator, while no analogous mechanism is present in linear FE.
This paper and \citet{HalkiewiczCycle2026} form a shared program of adequacy
diagnostics for saturated fixed-effect designs: the companion treats
concentrated identifying variation with exact nuisance-annihilating contrasts,
this paper treats noisy continuous regressors, and the same implementation also
screens heterogeneous TWFE effects. The point is diagnostic separation---small
within variation alone is not weak identification, but concentrated scores and
measurement error are distinct failures requiring distinct remedies.
The conjecture becomes true once we introduce a bias source. We study the canonical one: classical measurement error in the treatment. Observed $X_i^* = X_i + \nu_i$ replaces the true $X_i$, with $\nu_i$ a mean-zero noise of variance $\sigma_\nu^2$ independent of $(X_i, u_i)$. Under Condition~(B), FE saturation asymptotically absorbs the same fraction $\rho$ of the noise variance and the true signal variance (exactly so under (B1)), leaving the within reliability unaffected by $\rho$ and the residual identifying variation smaller by the common factor $(1-\rho)$. We work under the local-to-zero drift $\sigma_\nu^2 = c^2/n$, which keeps the noise commensurable with the residual treatment variance, and show that the FE-OLS $t$-statistic converges to a non-central normal:
\[
T_n^* \Rightarrow N(\eta, 1), \qquad
\eta = -\frac{\beta_0\, c^2\sqrt{1-\rho}}{\sigma \sqrt{\tau^2 + c^2}}.
\]
The non-centrality is non-zero precisely when $c > 0$, and decreasing in $\tau^2$, $\sigma$, and (at fixed $\tau^2$) in $\rho$ through the overall $\sqrt{1-\rho}$ scaling of the design's total identifying variation, \emph{not} through any differential absorption of noise relative to signal. The two-sided size is \emph{even} in $\eta$, so its leading distortion is quadratic, $z_{1-\alpha/2}\phi(z_{1-\alpha/2})\,\eta^2$; inverting it at a tolerance $\delta$ yields a closed-form Stock--Yogo critical value:
\[
\tau^2_{\mathrm{crit}}(\rho, c^2, \beta_0, \sigma; \alpha, \delta)
= \frac{\beta_0^2\, c^4 (1-\rho)\, z_{1-\alpha/2}\,\phi(z_{1-\alpha/2})}{\sigma^2\delta} - c^2.
\]
The operational threshold is zero when this algebraic boundary is negative.
\begin{framed}
\noindent\textbf{What the diagnostic certifies.} The non-centrality is proportional to $\beta_0$: classical attenuation shrinks a zero to a zero, so under $H_0:\beta_0=0$ there is no size distortion, and the significance test of \emph{no effect} is never the object at risk. What measurement error distorts is inference about the \emph{magnitude}: the $t$-test of $H_0:\beta=\beta_0$ at the true nonzero $\beta_0$ --- equivalently, the coverage of the conventional confidence interval, which is centered near the attenuated $\lambda\beta_0$ and misses the truth at the non-central rate. Throughout the paper, the object the diagnostic speaks to is confidence-interval coverage of $\beta_0$ --- the credibility of the reported magnitude, not the significance verdict; Definition~\ref{def-breakdown} encodes this by setting the breakdown reliability to zero when $\beta_0=0$. That definition also fixes the paper's two verdict words, which are never used interchangeably: a \emph{point pass} is descriptive, a \emph{certificate} is conservative and accounts for the pilot's sampling error.
\end{framed}
A second question the reader should ask up front: if the reliability $\lambda$ is knowable, why diagnose rather than simply correct --- report $\widehat\beta^*/\widehat\lambda$ with standard errors inflated by $1/\widehat\lambda$ and dispense with thresholds? The answer is the Stock--Yogo logic itself: \emph{the correction requires a point estimate of $\lambda$; the diagnostic requires only a lower bound}. The verdict ``conventional inference is adequate provided reliability exceeds the breakdown value $\lambda^{\dagger}$, a bound the validation literature comfortably supports'' survives any imprecision in the noise pilot that keeps it above $\lambda^{\dagger}$, whereas the corrected estimator inherits the pilot's error one-for-one, together with its sampling noise (Corollary~\ref{cor-slope}). Certification is the robust use of weak side information; correction is the fragile use of strong side information.
The structure of this threshold differs from the IV case in three ways. First, it is two-dimensional in the regime parameter: $\tau^2$ enters through the standardization $\sqrt{\tau^2+c^2}$ exactly as in a signal-to-noise ratio, and $\rho$ enters separately, through the overall $\sqrt{1-\rho}$ discount common to signal and noise. Heavier FE saturation \emph{lowers} the threshold, but not because it absorbs noise preferentially: it shrinks the design's total identifying variation, which raises the standard error at any fixed $\tau^2$. Second, the threshold depends on $|\beta_0|$ and $\sigma$ (the parameter being estimated and the residual standard deviation), so unlike the unit-free Stock--Yogo IV thresholds, this one is data-dependent. Third, the threshold requires an external estimate of the measurement-error variance $c^2$. This estimate typically comes from test-retest validation, published reliability ratios, or, for expert-coded indices, the measurement model's own posterior uncertainty; the diagnostic cannot be computed from the regression output alone.
\subsection{Relation to Griliches--Hausman} The interaction between panel transformations and measurement error is classical: \citet{GH} showed that within and difference transformations \emph{amplify} attenuation bias, because they remove signal variance while (largely) preserving noise variance, and they proposed estimators (contrasts across differencing lengths, and instrumenting with lags) that exploit the differential amplification to identify and remove the bias. The present paper is complementary and answers a different question. \citet{GH} treat the problem as one of \emph{estimation}: given that the bias exists, construct a consistent estimator. We treat it as one of \emph{inference adequacy}: given that the practitioner will run conventional FE-OLS on the observed regressor, characterize exactly when the resulting $t$-test remains size-controlled, and supply the critical value that certifies it. The two-dimensional $(\rho, \tau^2)$ threshold, the separate roles of the within reliability $\lambda$ and the FE dimension $\rho$ in it, and the local-drift non-centrality analysis are, to our knowledge, new. When a specification fails the threshold, the Griliches--Hausman estimators (or a fixed-effect Anderson--Rubin test, which we develop in companion work in progress) are the natural remedies; the diagnostic tells the practitioner whether a remedy is needed at all.
\subsection{Classical error in a continuous regressor only}\label{sec-scope}
One restriction is important enough to state before the results rather than after them, because violating it inverts the diagnostic's conclusion rather than merely weakening it. Everything in this paper assumes \emph{classical} measurement error: $X^*_i = X_i + \nu_i$ with $\nu_i$ mean-zero and independent of $(X_i, u_i)$, formalized in Assumption~\ref{ass-nu}. The entire apparatus---the attenuation factor $\lambda$, the sign of the non-centrality, the breakdown reliability, the direction in which the naive pilot errs---rests on that independence.
It fails, by construction and not by approximation, for a mismeasured \emph{binary} treatment. If $X_i \in \{0,1\}$ then the misclassification error $\nu_i = X^*_i - X_i$ can only be $+1$ when $X_i = 0$ and $-1$ when $X_i = 1$, so $\operatorname{Cov}(\nu_i, X_i) < 0$ necessarily: the error is nonclassical, mean-independence fails, and the attenuation factor is no longer the reliability ratio. This is not a new observation --- \citet{Aigner1973} derived the attenuation for a mismeasured binary regressor and showed it does not take the classical form, and \citet{Bollinger1996} shows that what is identified under misclassification is a \emph{bound} on the coefficient rather than a point-corrected estimate --- which is precisely why a reliability-based threshold has nothing to certify there. Since binary treatments---union status, program participation, treatment take-up---are among the most commonly mismeasured regressors in fixed-effect panels \citep[see][on the union case]{Card1996}, we state the consequence plainly: \textbf{applying the diagnostic of this paper to a binary treatment is a misuse, and its verdict carries no warrant.} A reliability ratio computed for such a regressor does not index the object $\lambda$ that appears in Definition~\ref{def-breakdown}, and the resulting pass or flag is not interpretable. The applications below use continuous regressors (expert-coded democracy indices, the log wage) throughout, and step~0 of the protocol in Section~\ref{sec-protocol} makes the check explicit. Extending the derivation pattern to nonclassical misclassification requires its own non-centrality calculation and is left to future work (Section~\ref{sec-ext}).
\subsection{``Is this not just bias divided by standard error?''}\label{sec-notjust}
The non-centrality $\eta$ is, algebraically, an attenuation bias divided by a standard error. Three things make it an econometric object rather than a ratio, and each is a place where the informal calculation gives the wrong answer.
\emph{First, the ratio is not a verdict until it has a limit experiment attached.} It becomes a statement about size only once one knows the law of the $t$-statistic under contamination; under the local drift that law is $N(\eta,1)$ (Theorem~\ref{thm-noncentral}). Without it, the natural conjecture --- that attenuation of the point estimate translates one-for-one into over-rejection --- is simply wrong. The two-sided size is \emph{even} in $\eta$, so the leading distortion is quadratic and the critical value is not linear in the ratio; reasoning informally from ``bias over standard error'' misprices the threshold by an order in $\eta$.
\emph{Second, correcting and certifying need different information.} Correction requires a point estimate of $\lambda$ and inherits its error one-for-one; certification requires only a \emph{lower bound}, because the verdict is monotone in $\lambda$. This is the Stock--Yogo logic, and it is what makes the exercise feasible with the side information validation studies actually supply --- reliability ranges, not reliability point estimates. A bare ratio cannot exploit the asymmetry, having no threshold to be monotone against.
\emph{Third, the obvious way to compute the ratio is anti-conservative.} Plugging in the estimated coefficient uses the \emph{attenuated} $\widehat\beta^*_K$, understating $|\eta|$ by the factor $\lambda$ and passing specifications that should fail (Proposition~\ref{prop-pilot}); Design~3 exhibits one whose true size is $16\%$. The error runs in the direction users care about, and removing it requires the reliability correction --- exactly the step a bookkeeping account omits.
Three further properties follow from the analysis rather than motivating it, and are developed where they arise: reliability and saturation are distinct primitives entering through different channels (Remark~\ref{rem-rho}); the breakdown reliability is computable from a published $t$-statistic alone, with no data and no noise pilot (Definition~\ref{def-breakdown}); and the verdict depends on which variance estimator the applied author actually reported, to the point of reversal (Theorem~\ref{thm-cluster}).
Our contributions are:
\begin{enumerate}[label=(\roman*)]
\item Formal demonstration that no $\tau^2$-driven size threshold arises in the baseline FE-OLS model (Theorem~\ref{thm-no-distortion}): FE saturation alone is not a weak-identification problem.
\item Derivation of the non-centrality parameter under EIV with local drift (Theorem~\ref{thm-noncentral}), with the FE dimension $\rho$ entering only through a common $\sqrt{1-\rho}$ discount of signal and noise alike (not through differential absorption), and a local power result (Proposition~\ref{prop-power}): power is attenuated by the factor $\sqrt{(1-\rho)\lambda}$ and shifted by $\eta$.
\item A feasible representation of the non-centrality, $|\eta| = (|\beta_0|/\sigma)(1-\lambda)\sqrt{\tau^{*2}}$ (Corollary~\ref{cor-feasible}), in which every quantity except one external reliability pilot is regression output; and a closed-form critical-value formula (Corollary~\ref{cor-cv}).
\item A reliability-corrected pilot for $\beta_0$ (Proposition~\ref{prop-pilot}). Plugging the attenuated point estimate $\widehat\beta^*_K$ into the threshold understates $|\beta_0|$ by the factor $\lambda$ and makes the diagnostic \emph{anti-conservative}; the corrected pilot removes this.
\item The cluster-robust theory (Section~\ref{sec-cluster}): a cluster-level score CLT (Lemma~\ref{lem-cluster-clt}), consistency of the Arellano variance estimator in the many-fixed-effect regime under a checkable projection-compatibility condition (Lemmas~\ref{lem-nest}--\ref{lem-crve}), and the resulting non-centrality $\eta_{CR}=\eta/\sqrt\psi$ (Theorem~\ref{thm-cluster}). Separately, a \emph{formal certification theorem} (Proposition~\ref{prop-certificate}): evaluating the diagnostic at a $(1-\gamma)$ upper confidence bound for $|\beta_0|$ controls the false-certification probability at $\gamma$, which is what distinguishes a certificate from the descriptive point pass (Remark~\ref{rem-plugin}) and what the paper's title refers to. Two by-products are of independent practical interest: the conventional $\tfrac{n-1}{n-K_n}$ small-sample factor must be \emph{omitted} from the CRVE in saturated designs, and the feasible rescaling remains valid even though its two ingredients are separately inconsistent (Corollary~\ref{cor-cluster-feasible}).
\item A simulation design that calibrates the accuracy of the local-drift approximation, including the comparison against fixed finite-$n$ noise (which delimits when the approximation is trustworthy), and an empirical application in which the measurement-error variance is externally observable.
\end{enumerate}
Companion work in progress develops a complementary tool: a fixed-effect Anderson--Rubin test \citep[after][]{AndersonRubin} with uniform validity over $\tau^2 \in (0, \infty]$, providing valid inference precisely for the specifications that fail this paper's threshold.
\subsection{Related literature} Classical measurement error in panel regressions goes back to \citet{GH}; \citet{BBM} survey validation-study evidence on the magnitude of survey measurement error, and \citet{Schennach2016} surveys the modern identification-based literature. \citet{BK} provide the canonical external reliability estimates for earnings data (CPS matched to Social Security records), and \citet{BoundEtAl94} the PSID validation-study estimates, including the sharp drop in reliability after within/difference transformations, the empirical fact that motivates our drift calibration in Section~\ref{sec-drift}. For expert-coded political indices, \citet{PMM} estimate an explicit measurement model whose posterior dispersion supplies a direct observation-level noise estimate; this is the basis of our lead application. The weak-identification literature we borrow from is \citet{SS, SY, Moreira2003, OleaPflueger2013, ASS, LMMP}. \citet{CJN} and \citet{JV} develop the high-dimensional FE asymptotics in which $\rho > 0$ (with $\tau^2 = \infty$ in our parametrization); we work within that framework but with the additional dimension $\tau^2 < \infty$. \citet{KSS} treat the AKM bias-correction problem with leave-one-out techniques, and \citet{Jochmans2022} develops heteroskedasticity-robust inference under many covariates; together with \citet{CJN} these supply the variance-estimation side of the many-regressor asymptotics we work in. The two auxiliary results this paper itself needs --- a self-normalized score CLT and a Hessian concentration lemma under the $(\rho,\tau^2)$ drift --- are stated and proved in \ref{app-aux}, keeping the paper self-contained. To our knowledge, no formal Stock--Yogo-style critical values for FE saturation under measurement error have been derived.
Sections~\ref{sec-setup}--\ref{sec-baseline} develop the setup and the baseline no-threshold result; Section~\ref{sec-eiv} introduces the EIV model and derives the non-centrality; Section~\ref{sec-cv} gives the critical-value formula, feasible diagnostic, and protocol; Sections~\ref{sec-sim}--\ref{sec-app} report the simulations and applications; Sections~\ref{sec-ext}--\ref{sec-conc} record extensions and conclude.
\section{Setup}\label{sec-setup}
Let $\{(Y_i, X_i)\}_{i=1}^n$ be a sample from a population on $(\Omega, \mathcal F, \mathbb P)$, with $Y_i \in \mathbb R$ and $X_i \in \mathbb R$ a scalar treatment. The structural model is
\begin{equation}\label{eq-model}
Y_i = X_i \beta_0 + u_i, \qquad \operatorname{\mathbb E}[u_i \mid X_i, \mathcal G_\infty] = 0,
\end{equation}
where $\mathcal G_\infty = \sigma(\bigcup_{K \geq 1} \mathcal G_K)$ is the limit of an increasing sequence of sub-$\sigma$-algebras representing successive fixed-effect specifications. Let $D_K$ be the $n \times d_K$ matrix of FE dummies, $P_K$ the orthogonal projection onto the FE column space, and $M_K = I_n - P_K$ the FE annihilator; all partialling-out identities used below are instances of the Frisch--Waugh--Lovell theorem \citep{FrischWaugh,Lovell1963}, of which \citet{LovellFWL} gives a compact modern proof. Write $\widetilde X_K = M_K X$. Throughout, the fitted dummies \emph{saturate} the specification: $D_K$ spans the $\mathcal G_K$-measurable functions of the cell labels, so that $\operatorname{\mathbb E}[X \mid \mathcal G_K] \in \mathrm{col}(D_K)$ and hence $M_K\,\operatorname{\mathbb E}[X\mid\mathcal G_K]=0$ (and likewise any additive $\mathcal G_K$-measurable component of the outcome is annihilated). This is the assumption under which $X'M_KX$ equals the pure-deviation quadratic form $\xi'M_K\xi$, $\xi := X - \operatorname{\mathbb E}[X\mid\mathcal G_K]$, used in Lemma~\ref{lem-clt} below.
The FE-residualized OLS estimator (computed in this section from the \emph{true} $X$) is $\widehat\beta_K = (X' M_K Y) / (X' M_K X)$. The population residual treatment variance is
\[
Q_K = \operatorname{\mathbb E}[(X - \operatorname{\mathbb E}[X \mid \mathcal G_K])^2] = \operatorname{\mathbb E}[\operatorname{Var}(X \mid \mathcal G_K)],
\]
with sample analogue $\widehat Q_K = n^{-1} X' M_K X$. The asymptotic regime is governed by
\begin{equation}\label{eq-drift}
\rho_n = \frac{d_{K_n}}{n} \to \rho \in [0, 1), \qquad \tau_n^2 = nQ_{K_n} \to \tau^2 \in (0, \infty].
\end{equation}
The case $\tau^2 = \infty$ recovers the many-covariates regime of \citet{CJN} and \citet{JV}; the case $\rho = 0$, $\tau^2 = \infty$ is textbook FE asymptotics. The non-centrality \emph{formula} nests these limits by continuity (Remark~\ref{rem-nesting}); the theorems below are proved for the interior $0 < \tau^2 < \infty$, $\rho \in [0,\bar\rho]$ with $\bar\rho < 1$, and the boundaries $\tau^2 \in \{0,\infty\}$ and $\rho \to 1$ are separate limiting regimes at which the proofs' leverage and concentration constants are not uniform.
The joint CLT under this drift is the foundational result on which this paper builds. We state it as Lemma~\ref{lem-clt} and prove it in full in \ref{app-aux}, so that the paper is self-contained.
\begin{assumption}[Regularity]\label{ass-reg}\hfill
\begin{enumerate}[label=(\roman*)]
\item $\{(X_i, u_i, \mathcal G_K)\}$ is i.i.d.\ across $i$ for each $K$, with finite eighth moments of $X$ and $u$.
\item $\rho_n \to \rho \in [0, 1)$, $\tau_n^2 \to \tau^2 \in (0, \infty]$.
\item Leverage control: there is $\kappa > 0$ with $\max_i (P_{K_n})_{ii} \le 1 - \kappa$ with probability approaching one, and $\max_i \widetilde X_{K_n, i}^2 / (nQ_{K_n}) = o_p(1)$. (A vanishing $\max_i (P_{K_n})_{ii}$ would force $\rho = 0$, since $\max_i (P_{K_n})_{ii} \ge \operatorname{tr}(P_{K_n})/n = \rho_n$; the many-fixed-effect regime $\rho > 0$ requires only that no single cell dominates, i.e.\ leverage bounded away from one. Only the treatment-leverage half is used in the score CLT below.)
\item Conditional variance: there exists $\omega^2 \in (0, \infty)$ with
\[
(nQ_{K_n})^{-1} \sum_i \widetilde X_{K_n, i}^2 \sigma^2(X_i, \mathcal G_{K_n, i}) \to_p (1-\rho)\,\omega^2,
\]
so that $\omega^2$ is the design-free (fully-diffuse) conditional-variance limit and the observable score variance carries the same $(1-\rho)$ discount as the Hessian limit in Lemma~\ref{lem-clt}; under conditional homoskedasticity $\omega^2 = \sigma^2$.
\item Lindeberg condition: for every $\varepsilon > 0$,
\[
\resizebox{0.97\linewidth}{!}{$\displaystyle (nQ_{K_n})^{-1} \sum_i \operatorname{\mathbb E}\big[\widetilde X_{K_n, i}^2 u_i^2 \, \mathbf 1\{|\widetilde X_{K_n, i} u_i| > \varepsilon \sqrt{nQ_{K_n}}\}\big] \to 0$}.
\]
\item Uniform conditional moment: for some $\epsilon > 0$,
$\sup_{i,n} \operatorname{\mathbb E}[\,|u_i|^{2+\epsilon} \mid X_i,\mathcal G_{K_n}\,]
\le C_u < \infty$. (The conditioning includes the regressor because the score
proof conditions on $\mathcal F_n=\sigma(X,D_{K_n})$; a bound conditional only on
$\mathcal G_{K_n}$ would not in general survive further conditioning on $X_i$.
Together with a maximal-weight condition this yields the Lyapunov, hence
Lindeberg, condition for \emph{independent but not necessarily identically
distributed} weighted sums of the $u_i$ --- the form used for the contaminated
score in Theorem~\ref{thm-noncentral}; it is implied by the finite eighth moment
of (i) when the conditional law of $u_i$ does not degenerate across cells and
regressor values.)
\end{enumerate}
\end{assumption}
\begin{lemma}[Joint CLT under the $(\rho,\tau^2)$ drift]\label{lem-clt}
Under Assumption~\ref{ass-reg} and the balance conditions (B)--(C) of \ref{app-aux},
\[
\begin{pmatrix}
\dfrac{X' M_{K_n} u}{\sqrt{nQ_{K_n}}} \\[4pt]
\dfrac{X' M_{K_n} X}{nQ_{K_n}} \\[4pt]
\dfrac{u' M_{K_n} u}{n - d_{K_n}}
\end{pmatrix}
\Rightarrow
\begin{pmatrix} \mathcal Z \\ 1-\rho \\ \sigma^2 \end{pmatrix},
\qquad \mathcal Z \sim N(0, (1-\rho)\,\omega^2),
\]
with the second and third coordinates converging in probability. Under conditional homoskedasticity, $\omega^2 = \sigma^2$, so the score variance is $(1-\rho)\sigma^2$, not $\sigma^2$.
\end{lemma}
\emph{Why the Hessian limit is $1-\rho$.} The second coordinate is worth a word, since it drives the critical-value formula. With $\xi_i := X_i - \operatorname{\mathbb E}[X_i\mid\mathcal G_{K_n}]$ and $X'M_{K_n}X=\xi'M_{K_n}\xi$ under saturation, the conditional mean is $\sum_i(1-h_{ii})q_i$, where $q_i=\operatorname{Var}(\xi_i\mid\mathcal D_n)$ and $h_{ii}=(P_{K_n})_{ii}$. Under (B1) this equals $Q_{K_n}\operatorname{tr}(M_{K_n})=(1-\rho_n)nQ_{K_n}$ exactly; under (B2) leverage balance gives the same expression up to $1+o_p(1)$. Saturation alone does not imply the common trace discount. The exact trace identity does apply to homoskedastic errors and to the i.i.d.\ noise form $\nu'M_{K_n}\nu$, while Condition~(B) supplies it for the signal. The proof in \ref{app-aux} makes these distinctions explicit and conditions on $\mathcal F_n=\sigma(X,D_{K_n})$ for the score CLT.
\section{No Baseline Threshold}\label{sec-baseline}
A natural conjecture, building on the analogy with weak instruments, is that small $Q_K$ produces a Stock--Yogo problem requiring its own critical-value table. We first show this conjecture is incorrect in the baseline model.
Let $\widehat u = M_{K_n}(Y - X\widehat\beta_{K_n})$ be the regression residual after removing \emph{both} the fixed effects and the treatment (equivalently, the residual from the full OLS of $Y$ on $[D_{K_n}, X]$; by the Frisch--Waugh--Lovell theorem \citep{FrischWaugh, Lovell1963, LovellFWL}, $\widehat\beta_{K_n}$ is the coefficient on $X$ there), and let $\widehat\sigma^2_{\mathrm{CJN}} = \widehat u'\widehat u / (n - d_{K_n} - 1)$ be the associated residual variance estimator, with
\[
T_n^{\mathrm{CJN}} = \frac{\widehat\beta_{K_n} - \beta_0}{\widehat\sigma_{\mathrm{CJN}} / \sqrt{X' M_{K_n} X}}.
\]
\begin{theorem}[No $\tau^2$-driven size distortion in the baseline model]\label{thm-no-distortion}
Under Assumption~\ref{ass-reg}, suppose $u\mid(D_{K_n},X)\sim N(0,\sigma^2I_n)$ and $X' M_{K_n} X > 0$. Then $T_n^{\mathrm{CJN}} \sim t_{n - d_{K_n} - 1}$ exactly conditional on $(D_{K_n},X)$. The distortion relative to $N(0,1)$ satisfies
\[
\mathbb P(|T_n^{\mathrm{CJN}}| > z_{1-\alpha/2}) - \alpha = \frac{(z_{1-\alpha/2}^3 + z_{1-\alpha/2}) \phi(z_{1-\alpha/2})}{2(n - d_{K_n} - 1)} + O(n^{-2}),
\]
which is $O(1/n)$ and does not depend on $\tau_n^2$.
\end{theorem}
\begin{proof}[Proof sketch]
Conditional on $(D_{K_n}, X)$, the numerator $\widehat\beta_{K_n}-\beta_0 = b'u$ (a linear form, $b = M_{K_n}X/(X'M_{K_n}X)$) and the residual sum of squares $\widehat u'\widehat u = u'M_{[K,X]}u$ (a quadratic form in the augmented annihilator $M_{[K,X]}$) are independent for Gaussian $u$, because $M_{[K,X]}M_{K_n}X = 0$. This is exactly why the residual variance must be formed after projecting out $X$: with $M_{K_n}$ in place of $M_{[K,X]}$ independence fails. Hence $T_n^{\mathrm{CJN}} \sim t_{n-d_{K_n}-1}$ exactly, and the Fisher--Cornish expansion \citep[26.7.8]{AS72} gives the displayed $O(1/n)$ distortion, uniform in $\rho \in [0,\bar\rho]$ and independent of $\tau_n^2$. Full details in \ref{app-nodist}.
\end{proof}
\begin{remark}[Why the IV analogy breaks]\label{rem-no-IV}
TSLS is biased of order $1/\mu^2$ because substituting $\widehat X$ for $X$ multiplies the first-stage residual into the second-stage residual, producing a term with expectation proportional to $\operatorname{\mathbb E}[vu]$. FE-residualized OLS contains no such substitution: $M_KX$ is an exact projection onto a known subspace, so $\operatorname{\mathbb E}[X'M_Ku]=0$ in finite samples and the estimator is unbiased. Any bias source capable of generating a Stock--Yogo threshold must therefore come from outside the baseline model.
\end{remark}
\begin{remark}[Heteroskedasticity]\label{rem-het}
Theorem~\ref{thm-no-distortion} is stated under homoskedasticity, where the exact $t$ distribution is available. Under heteroskedasticity, leave-one-out (HC2-type) variance estimators restore asymptotic validity of the baseline $t$-test in many-covariate designs \citep{CJN, KSS, Jochmans2022}; the ``no $\tau^2$ threshold'' conclusion is unchanged, though the exact finite-sample distribution is lost.
\end{remark}
We now introduce a natural bias source: classical measurement error in the treatment.
\section{The Non-Centrality Theorem}\label{sec-eiv}
\subsection{Model and drift}\label{sec-eiv-model}
Suppose the analyst observes
\begin{equation}\label{eq-eiv}
X_i^* = X_i + \nu_i,
\end{equation}
where $X_i$ is the true treatment (unobserved), and $\nu_i$ is measurement error with $\operatorname{\mathbb E}[\nu_i \mid \mathcal G_{K_n}] = 0$, $\operatorname{Var}(\nu_i \mid \mathcal G_{K_n}) = \sigma_\nu^2$, and $\nu_i \perp (X_i, u_i)$ conditional on $\mathcal G_{K_n}$. We retain $Q_{K_n}$ for the residual variance of the \emph{true} treatment. The following assumption records the additional moment and tail conditions on the noise used in the variance calculations of Lemma~\ref{lem-atten} and the leverage argument of Lemma~\ref{lem-lev-star}.
\begin{assumption}[Measurement-error regularity]\label{ass-nu}\hfill
\begin{enumerate}[label=(\roman*)]
\item \emph{(Independence.)} $\{\nu_i\}_{i=1}^n$ are independent conditional on $\mathcal G_{K_n}$, with $\operatorname{\mathbb E}[\nu_i\mid\mathcal G_{K_n}]=0$, $\operatorname{Var}(\nu_i\mid\mathcal G_{K_n})=\sigma_{\nu,n}^2$, and $\nu\perp(X,u)\mid\mathcal G_{K_n}$.
\item \emph{(Scaled moment bound.)} For some fixed $r>1$, the standardized errors have a uniformly bounded $2r$-th moment: $\operatorname{\mathbb E}[\,|\nu_i/\sigma_{\nu,n}|^{2r}\mid\mathcal G_{K_n}\,]\le C_\nu<\infty$ for all $i,n$ --- equivalently $\operatorname{\mathbb E}[\,|\nu_i|^{2r}\mid\mathcal G_{K_n}\,]\le C_\nu\,\sigma_{\nu,n}^{2r}$, the natural requirement that the noise has no heavier tails than a rescaled fixed distribution (it is not a bare finite-moment bound: the $2r$-th moment must scale as the $2r$-th power of the standard deviation, which is what the leverage argument uses). The canonical case is a uniformly bounded standardized fourth moment ($r=2$), which is also what the quadratic-form variance bound of Lemma~\ref{lem-atten} uses; the fourth moments need not be identical across observations.
\end{enumerate}
\end{assumption}
Assumption~\ref{ass-nu} imposes no tail condition beyond a finite fourth moment. Two moment orders appear, and we keep them separate. The quadratic-form variance in Lemma~\ref{lem-atten} --- and hence the denominator concentration used throughout, including the leverage \emph{ratio} of Lemma~\ref{lem-lev-star} and Theorem~\ref{thm-noncentral} --- uses $r=2$; the \emph{maximal-noise} bound of Lemma~\ref{lem-lev-star} (its first display) and Remark~\ref{rem-lev-moment} need only $r>1$. We therefore take $r=2$ as the standing requirement for the EIV theorems and flag the two places where $r>1$ alone suffices. This is the routine light-tail requirement of the measurement-error literature --- weaker than the sub-Gaussianity one might expect a maximal-leverage argument to need; sub-Gaussian tails only sharpen the rate (Remark~\ref{rem-lev-moment}) and are not required.
The analyst runs FE-OLS on $(Y, X^*)$, obtaining
\begin{equation}\label{eq-est-star}
\widehat\beta_K^* = \frac{X^{*\prime} M_K Y}{X^{*\prime} M_K X^*}.
\end{equation}
We adopt the \emph{local measurement-error drift}:
\begin{equation}\label{eq-local-drift}
\sigma_{\nu,n}^2 = \frac{c^2}{n}
\end{equation}
for a fixed scalar $c^2 \geq 0$.
\begin{lemma}[Attenuation under the drift]\label{lem-atten}
Under Assumption~\ref{ass-reg}, Conditions~(B)--(C) of \ref{app-aux}, Assumption~\ref{ass-nu} with $r = 2$ (uniformly bounded standardized fourth moments), and the drift~\eqref{eq-local-drift}:
\begin{enumerate}[label=(\alph*)]
\item (Weak information, $\tau^2 < \infty$.) The true and contaminated Hessians and the attenuation factor satisfy
\[
\begin{gathered}
X'M_{K_n}X \to_p (1-\rho)\tau^2, \qquad
X^{*\prime} M_{K_n} X^* \to_p \tau^{*2} := (1-\rho)(\tau^2 + c^2), \\[4pt]
\lambda := \frac{X'M_{K_n}X}{X^{*\prime}M_{K_n}X^*} \to_p \frac{\tau^2}{\tau^2+c^2} \in (0, 1].
\end{gathered}
\]
Under Condition~(B), the signal and measurement noise are discounted by the \emph{same asymptotic} factor $(1-\rho)$ (exactly in conditional mean under (B1)), so $\lambda$, the within reliability, does not depend on $\rho$: saturation shrinks the design's total identifying variation but does not change the limiting relative share of signal to noise within it. The slope estimator is \emph{not} consistent; its limiting law --- centered at $\beta_0\lambda$ with nondegenerate Gaussian error --- is recorded in Corollary~\ref{cor-slope}, once the score CLT of Theorem~\ref{thm-noncentral} is available.
\item (Strong information, $nQ_{K_n} \to \infty$.) The slope estimator is consistent for the (now asymptotically unattenuated) value, $\widehat\beta_K^* \to_p \beta_0$, since $\lambda \to 1$.
\end{enumerate}
\end{lemma}
\begin{proof}
Deferred to \ref{app-atten}. In the weak-information regime, Lemma~\ref{lem-clt} gives $X'M_{K_n}X\to_p(1-\rho)\tau^2$. Expanding $X^{*\prime}M_{K_n}X^*$, the cross term $X'M_{K_n}\nu$ is $o_p(1)$, while the bounded-kurtosis quadratic-form bound of Lemma~\ref{lem-qf} gives $\nu'M_{K_n}\nu\to_p c^2(1-\rho)$ --- exactly the same discount, applied to the noise. In the strong-information regime the same cross and noise terms are negligible relative to $nQ_{K_n}$; the estimator claim then follows by comparing the score with the denominator.
\end{proof}
The finite-$n$ attenuation factor is
\begin{equation}\label{eq-lambda}
\lambda_n := \frac{Q_{K_n}}{Q_{K_n} + \sigma_{\nu, n}^2} = \frac{\tau_n^2}{\tau_n^2 + c^2} \to \lambda,
\end{equation}
which is exactly the population conditional-within reliability of the observed regressor. Under Condition~(B) it is also the asymptotic share of sample post-projection variation in $X^*$ that is signal (and under (B1) the signal and noise conditional means have the same finite-sample trace factor). Thus no $\rho_n$ appears in this ratio; the common asymptotic factor survives only in the total scale $\tau_n^{*2}=(1-\rho_n)(\tau_n^2+n\sigma_{\nu,n}^2)$. Everything in this paper is organized around $\lambda$.
\subsection{Interpreting the drift}\label{sec-drift}
The drift $\sigma_\nu^2 = c^2/n$ needs justifying, because it is doing real work: it is the unique rate at which the projected noise $c^2(1-\rho)$ and the residual signal $\tau^2$ remain of the same order, and hence the unique rate at which the limit experiment produces a non-degenerate threshold. Fixed $\sigma_\nu^2 > 0$ makes attenuation $O(1)$ and the size distortion total; $\sigma_\nu^2 = o(1/n)$ makes measurement error vanish from the limit entirely. Three points establish that the knife-edge is the empirically relevant calibration rather than a mathematical convenience.
First, the drift is an \emph{approximation device with a finite-sample target}, in precisely the sense of \citet{SS}: weak-instrument asymptotics model the first stage as local to zero not because first stages literally shrink with $n$, but because the resulting limit distribution approximates the finite-sample distribution well whenever the concentration parameter is moderate. The same logic applies here. For a balanced treatment design satisfying (B1), the finite-$n$ calibration is
\begin{equation}\label{eq-eta-n}
\eta_n \;=\; -\frac{\beta_0\, \sigma_\nu^2\, \sqrt{n - d_{K_n}}}{\sigma\, \sqrt{\, Q_{K_n} + \sigma_\nu^2\,}}
\;=\; -\frac{\beta_0}{\sigma}\,(1 - \lambda_n)\,\sqrt{(n-d_{K_n})(Q_{K_n}+\sigma_\nu^2)},
\end{equation}
every ingredient of which is a fixed-$n$ quantity: no drifting sequence appears. The noise identity $\operatorname{\mathbb E}[\nu'M_{K_n}\nu]=\sigma_\nu^2(n-d_{K_n})$ is exact; under (B1), $\operatorname{\mathbb E}[X'M_{K_n}X\mid\mathcal D_n]=Q_{K_n}(n-d_{K_n})$ is exact as well. Under (B2), equation~\eqref{eq-eta-n} is the corresponding first-order calibration because the signal trace equality is only asymptotic. The theorems below say that $N(\eta_n, 1)$ approximates the law of the $t$-statistic whenever $\eta_n$ is moderate, and the balanced simulations in Section~\ref{sec-sim} (Design 4) quantify the quality of that approximation \emph{under a fixed finite-$n$} $\sigma_\nu^2$ (a single constant per cell), which is the direct check that the device is an approximation and not an artifact.
Second, the regime $\eta_n$ moderate is where applied work actually lives when FE saturate. Rewrite the noise-to-signal ratio as
\[
\frac{c^2}{\tau_n^2} = \frac{\sigma_\nu^2}{Q_{K_n}} = \frac{1 - \lambda_n}{\lambda_n},
\]
a population conditional-within ratio in which the fixed-effect dimension $\rho_n$ does not appear. Under Condition~(B), it also equals the limiting sample post-projection noise-to-signal ratio.
The empirical question is therefore: is the \emph{within reliability} $\lambda_n$ of real saturated-FE regressors close enough to one that measurement error is negligible, or close enough to zero that attenuation is total, or in between, where the threshold binds? The validation-study literature answers: in between, and increasingly so as saturation rises. Cross-sectional reliability of survey earnings is high (on the order of $0.7$--$0.85$ in the CPS--Social Security match of \citealp{BK}); but reliability after differencing or within transformations drops sharply, to roughly $0.5$--$0.65$ in the PSID validation study \citep{BoundEtAl94}, precisely because the transformation removes persistent signal while retaining transitory noise (the \citet{GH} amplification). Within reliabilities of $0.5$--$0.8$ correspond to $(1-\lambda)/\lambda$ between $0.25$ and $1$, exactly the range in which $\eta_n$ is moderate and the local approximation is designed to operate. The lead application of Section~\ref{sec-app-vdem} confirms this directly on real data: the realized $(1-\widehat\lambda_n)/\widehat\lambda_n$ is $0.11$ for the passing aggregate polyarchy specification and $0.83$ for the flagged legislative-constraints specification. Both sit inside the moderate window $[0.05, 10]$, so the local-drift regime is the empirically relevant one at both poles of the diagnostic.
Third, the drift formalizes a genuine comparative statics, not just an approximation: as researchers saturate ($\rho_n \uparrow$, $Q_{K_n} \downarrow$), the signal shrinks \emph{toward} the noise floor while the noise itself is design-invariant. The sequence $\sigma_\nu^2 = c^2/n$, $nQ_K \to \tau^2$ is the stylized path of that empirical practice: it holds the noise-to-residual-signal ratio at the level the practitioner's own saturation choices have produced. In this sense the drift is not modeling the world shrinking; it is modeling the analyst saturating.
\begin{remark}[Nesting]\label{rem-nesting}
The limit experiment nests the familiar cases. $c = 0$ gives $\eta = 0$: no measurement error, no threshold, Theorem~\ref{thm-no-distortion} applies. $\tau^2 = \infty$ gives $\eta = 0$: with abundant residual signal, measurement error of order $1/n$ is asymptotically irrelevant and standard strong-identification inference is recovered, consistent with the textbook observation that classical EIV bias is proportional to the noise-to-signal ratio. $\rho = 0$ recovers the cross-sectional (no-FE) threshold exactly, since the common discount $\sqrt{1-\rho}$ is then $1$; more generally $\rho$ enters $\eta$ \emph{only} through this overall scale factor, not through the reliability $\lambda$, which is $\rho$-free by Lemma~\ref{lem-atten}. The interesting region is $\tau^2 < \infty$, $c > 0$: saturated designs with residual signal of the same order as residual noise.
\end{remark}
\subsection{Non-centrality}\label{sec-eiv-thm}
The asymptotic normality of the score in the theorem below rests on a leverage condition for the \emph{contaminated} regressor $X^*$. We isolate it as a lemma; it is the one place where the moment bound of Assumption~\ref{ass-nu}(ii) is used, and it goes through under that bound alone, with no tail condition beyond a finite $2r$-th moment ($r>1$) needed.
\begin{lemma}[Leverage condition for the contaminated regressor]\label{lem-lev-star}
Suppose Assumption~\ref{ass-reg} and Conditions~(B)--(C) of \ref{app-aux} hold for the true regressor $X$ (so $\max_i \widetilde X_{K_n,i}^2/(X'M_{K_n}X)\to_p 0$ and $X'M_{K_n}X\to_p(1-\rho)\tau^2$, Lemma~\ref{lem-clt}), Assumption~\ref{ass-nu} holds (with moment order $r>1$ for the maximal-noise bound below, and $r=2$ for the leverage-ratio conclusion, which invokes Lemma~\ref{lem-atten}), and the drift~\eqref{eq-local-drift} holds with $\tau^2<\infty$. Write $\widetilde X^{*}_{K_n,i}:=(M_{K_n}X^*)_i=\widetilde X_{K_n,i}+\widetilde\nu_{K_n,i}$ with $\widetilde\nu_{K_n}:=M_{K_n}\nu$. Then
\[
\max_{1\le i\le n}\ \widetilde\nu_{K_n,i}^2 = O_p\!\big(c^2\, n^{1/r-1}\big)=o_p(1),
\qquad
\max_{1\le i\le n}\ \frac{\widetilde X^{*2}_{K_n,i}}{X^{*\prime}M_{K_n}X^*}\to_p 0 .
\]
\end{lemma}
\begin{proof}
Deferred to \ref{app-lev}. Splitting $\widetilde\nu_{K_n,i}=\nu_i-(P_{K_n}\nu)_i$, a union bound handles the raw part and Rosenthal's inequality the projected part, each giving $O_p(c^2 n^{1/r-1})=o_p(1)$ under the scaled moment bound of Assumption~\ref{ass-nu}(ii); the leverage ratio then follows from $\widetilde X^{*2}_{K_n,i}\le 2\widetilde X_{K_n,i}^2+2\widetilde\nu_{K_n,i}^2$ and Lemma~\ref{lem-atten}.
\end{proof}
\begin{theorem}[Non-centrality under EIV with local drift]\label{thm-noncentral}
Let $T_n^{\mathrm{CJN}*}(\beta_0) = (\widehat\beta_K^* - \beta_0) / (\widehat\sigma^*_{\mathrm{CJN}} / \sqrt{X^{*\prime} M_{K_n} X^*})$ be the CJN $t$-statistic computed from the contaminated regressor $X^*$, with $\widehat\sigma^{*2}_{\mathrm{CJN}} = \widehat u^{*\prime}\widehat u^*/(n - d_{K_n} - 1)$ and $\widehat u^* = M_{K_n}(Y - X^*\widehat\beta^*_K)$. The degrees of freedom match the baseline estimator of Section~\ref{sec-baseline}: the residual is formed after projecting out both the fixed effects and $X^*$, so $d_{K_n}+1$ parameters have been fitted. Nothing asymptotic turns on the $-1$, but it is the convention the replication code uses and the one printed regression output reports. Under Assumption~\ref{ass-reg} and Conditions~(B)--(C) of \ref{app-aux} applied to the true treatment $X$, Assumption~\ref{ass-nu} on $\nu$ with $r = 2$, the local drift~\eqref{eq-local-drift}, conditional homoskedasticity with the error conditionally independent of the regressor, $u \perp X \mid \mathcal G_{K_n}$ (so that, with $u \perp \nu \mid \mathcal G_{K_n}$, conditioning on $X^*$ preserves the mean-zero, variance-$\sigma^2$ conditional law of each $u_i$; the $u_i$ remain independent but need not be identically distributed across cells, and the uniform conditional moment of Assumption~\ref{ass-reg}(vi) discharges the contaminated-score Lindeberg condition), $\tau^2 < \infty$, and $H_0: \beta = \beta_0$,
\[
T_n^{\mathrm{CJN}*}(\beta_0) \Rightarrow N(\eta, 1), \qquad
\eta = -\frac{\beta_0\, c^2\sqrt{1-\rho}}{\sigma \sqrt{\tau^2 + c^2}}.
\]
\end{theorem}
\begin{proof}[Proof sketch]
Write $M = M_{K_n}$ and $N_n := X^{*\prime} M (Y - X^*\beta_0)$, so that $T_n^{\mathrm{CJN}*}(\beta_0) = N_n/(\widehat\sigma^*_{\mathrm{CJN}}\sqrt{X^{*\prime}MX^*})$ and, by Lemma~\ref{lem-atten}, $X^{*\prime}MX^*\to_p\tau^{*2}=(1-\rho)(\tau^2+c^2)$. Under $H_0$,
\[
N_n = \underbrace{X'Mu + \nu'Mu}_{\text{stochastic}} \;-\; \beta_0\underbrace{X'M\nu}_{o_p(1)} \;-\; \beta_0\underbrace{\nu'M\nu}_{\to_p\,c^2(1-\rho)},
\]
where the last term is the bias and the cross term $X'M\nu$ vanishes (Lemma~\ref{lem-atten}). The stochastic part is the \emph{single} score $X^{*\prime}Mu=\sum_i\widetilde X^{*}_{K_n,i}u_i$; the leverage condition of Lemma~\ref{lem-lev-star} supplies the maximal-weight condition, and with the uniform conditional $(2+\epsilon)$ moment of Assumption~\ref{ass-reg}(vi) the Lyapunov (hence conditional Lindeberg--Feller) CLT gives $X^{*\prime}Mu/\sqrt{\tau^{*2}}\Rightarrow N(0,\sigma^2)$ (by Lemma~\ref{lem-clt}, the score variance normalized by $nQ_{K_n}$ is $(1-\rho)\sigma^2$ under homoskedasticity, and $\tau^{*2}$ already carries the same $(1-\rho)$ factor, so no separate rescaling is needed), so $N_n\Rightarrow N(-\beta_0 c^2(1-\rho),\sigma^2\tau^{*2})$. Treating $X^{*\prime}Mu$ as a single weighted sum of the independent $u_i$, rather than splitting it into $X'Mu$ and $\nu'Mu$, avoids any joint-normality claim between the two pieces. A parallel expansion gives $\widehat\sigma^{*2}_{\mathrm{CJN}}\to_p\sigma^2$, and Slutsky assembles
\[
T_n^{\mathrm{CJN}*}(\beta_0)\Rightarrow \frac{-\beta_0c^2(1-\rho)}{\sigma\sqrt{\tau^{*2}}} = \frac{-\beta_0c^2(1-\rho)}{\sigma\sqrt{(1-\rho)(\tau^2+c^2)}} = \frac{-\beta_0c^2\sqrt{1-\rho}}{\sigma\sqrt{\tau^2+c^2}} = \eta.
\]
The full four-step argument, including the Lyapunov verification and the variance-estimator step, is in \ref{app-noncentral}.
\end{proof}
\begin{corollary}[Limiting law of the slope estimator]\label{cor-slope}
Under the conditions of Theorem~\ref{thm-noncentral} (in particular conditional homoskedasticity), with $\lambda=\tau^2/(\tau^2+c^2)$ and $\tau^{*2}=(1-\rho)(\tau^2+c^2)$ as in Lemma~\ref{lem-atten},
\[
\widehat\beta_K^* \;\Rightarrow\; \beta_0\lambda + N\!\big(0,\ \sigma^2/\tau^{*2}\big).
\]
The estimator is asymptotically centered at the attenuated value $\beta_0\lambda$, the \emph{classical} attenuation factor (no $\rho$-dependence); when $\tau^2 < \infty$ it is not consistent, and consistency ($\widehat\beta_K^*\to_p\beta_0\lambda$) is recovered only under strong information $nQ_{K_n}\to\infty$. The sampling variance $\sigma^2/\tau^{*2}$ does carry the $\rho$-dependence, since heavier saturation shrinks the total identifying variation $\tau^{*2}$ at fixed $(\tau^2,c^2)$.
\end{corollary}
\begin{proof}
Deferred to \ref{app-slope}: write $\widehat\beta_K^* = (\beta_0\, X^{*\prime}MX + X^{*\prime}Mu)/X^{*\prime}MX^*$ and apply Lemma~\ref{lem-atten} together with Step~2 of Theorem~\ref{thm-noncentral} and Slutsky.
\end{proof}
\begin{remark}[Sub-Gaussian tails only sharpen the rate]\label{rem-lev-moment}
Lemma~\ref{lem-lev-star} needs nothing beyond the finite $2r$-th moment of Assumption~\ref{ass-nu}(ii): any $r>1$ delivers $\max_i\widetilde\nu_{K_n,i}^2=O_p(c^2n^{1/r-1})=o_p(1)$, which is all the theorem uses. Stronger tails only sharpen the rate --- sub-Gaussian noise replaces $n^{1/r}$ by $\log n$ --- so we state the condition in the moment form standard in the measurement-error literature.
\end{remark}
\begin{remark}[Role of $\rho$ in the non-centrality]\label{rem-rho}
It is tempting to read $\eta = -\beta_0 c^2\sqrt{1-\rho}/(\sigma\sqrt{\tau^2+c^2})$ as saying that saturation \emph{preferentially} absorbs measurement noise, so that adding fixed effects is a targeted remedy for attenuation. It is not. Under Condition~(B), the within reliability $\lambda=\tau^2/(\tau^2+c^2)$ does not depend on $\rho$ (Lemma~\ref{lem-atten}): the signal and noise Hessians have the same asymptotic degrees-of-freedom factor, with the signal statement relying on treatment balance. What $\rho$ does is shrink the total identifying variation $\tau^{*2}=(1-\rho)(\tau^2+c^2)$ uniformly, so the whole $\rho$-dependence of $\eta$ collapses to $(1-\rho)/\sqrt{1-\rho}=\sqrt{1-\rho}$: heavier saturation lowers $|\eta|$ only by making the estimator noisier overall, exactly as any further covariate would. At $|\beta_0|/\sigma=1$, $\tau^2=c^2=5$ (so $\lambda=0.5$ throughout), $|\eta|$ falls from $1.50$ at $\rho=0.1$ to $1.00$ at $\rho=0.6$, the ratio $\sqrt{0.4/0.9}=0.667$ being exactly the $\sqrt{1-\rho}$ channel.
This is consistent with \citet{GH}, whose amplification result concerns a different comparative static: they hold the absolute noise variance fixed and let a finer transformation shrink the \emph{persistent} component of the signal while leaving transitory noise intact, so $\lambda$ falls. That mechanism operates on $X$'s covariance structure, which Assumption~\ref{ass-reg} does not engage; $\rho$ enters the reliability in neither analysis.
\end{remark}
\subsection{Local power}\label{sec-power}
The same limit experiment delivers the power side of the diagnostic.
\begin{proposition}[Local power]\label{prop-power}
Under the conditions of Theorem~\ref{thm-noncentral}, but with the true coefficient set to the information-standardized alternative $\beta_n = \beta_0 + b/\sqrt{nQ_{K_n}}$ for fixed $b \in \mathbb R$ (when $\tau^2 < \infty$ this is a fixed, non-vanishing displacement $\beta_n - \beta_0 \to b/\tau$, the natural scaling since the total information $nQ_{K_n}$ stays finite; it reduces to a conventional vanishing local alternative only in the strong-information limit $\tau^2 = \infty$),
\[
T_n^{\mathrm{CJN}*}(\beta_0) \Rightarrow N\!\Big(\frac{b}{\sigma}\sqrt{(1-\rho)\lambda} \;+\; \eta,\ 1\Big),
\qquad \lambda = \frac{\tau^2}{\tau^2 + c^2}.
\]
Consequently, relative to the fully diffuse no-FE benchmark (slope $b/\sigma$), measurement error and FE saturation jointly attenuate the information-standardized power slope by the factor $\sqrt{(1-\rho)\lambda}$ and shift the power curve by $\eta$. Relative to the error-free case at the \emph{same} $\rho$, whose slope is $(b/\sigma)\sqrt{1-\rho}$, measurement error alone contributes the factor $\sqrt\lambda$. Thus reliability supplies the classical attenuation-of-power channel, while $\sqrt{1-\rho}$ is the overall-scale effect that also governs $\eta$ (Remark~\ref{rem-rho}).
\end{proposition}
\begin{proof}
Deferred to \ref{app-power}. Under $\beta_n$ the numerator of Theorem~\ref{thm-noncentral} gains a term converging to $(1-\rho)b\tau$ (Lemma~\ref{lem-clt}: $X^{*\prime}MX\to_p(1-\rho)\tau^2$, so the added term $X^{*\prime}MX\cdot b/\sqrt{nQ_{K_n}} \to_p (1-\rho)\tau^2 \cdot b/\tau = (1-\rho)b\tau$); dividing by $\sigma\sqrt{\tau^{*2}}=\sigma\sqrt{(1-\rho)(\tau^2+c^2)}$ gives mean contribution $(1-\rho)b\tau/(\sigma\sqrt{(1-\rho)(\tau^2+c^2)}) = b\sqrt{1-\rho}\,\tau/(\sigma\sqrt{\tau^2+c^2}) = (b/\sigma)\sqrt{(1-\rho)\lambda}$, using $\tau/\sqrt{\tau^2+c^2}=\sqrt\lambda$.
\end{proof}
\begin{remark}[Reading the power result]\label{rem-power}
Proposition~\ref{prop-power} says the threshold is not only about size. Even a specification that passes the size threshold pays a power tax of $\sqrt{(1-\rho)\lambda}$; at within reliability $\lambda = 0.6$ and modest saturation $\rho=0.1$, the local slope falls to $\sqrt{0.9\times0.6}\approx74\%$ of its error-free value. Reporting $\widehat\lambda_n$ \emph{and} $\widehat\rho_n$ alongside the size diagnostic therefore serves double duty --- $\widehat\lambda_n$ summarizes the reliability tax, $\widehat\rho_n$ the saturation tax, and the two multiply rather than one subsuming the other. In the applications audited here $\widehat\rho_n$ is small ($0.03$ and $0.14$), so the $\sqrt{1-\rho}$ correction is minor in practice; it need not be for more heavily saturated designs (e.g.\ worker--firm networks), where reporting it separately matters more.
\end{remark}
\subsection{Cluster dependence}\label{sec-cluster}
Theorem~\ref{thm-noncentral} standardizes by the homoskedastic-i.i.d.\ variance, whereas applied panel practice clusters standard errors by unit \citep{Arellano1987, BDM}; see \citet{CameronMiller2015} and \citet{MNW2023} for current practice. Because the diagnostic is a bias-to-standard-error ratio, the choice of variance estimator moves it directly, and in the applications of Section~\ref{sec-app} it moves it enough to change a verdict. This subsection therefore develops the cluster-robust limit theory rather than asserting a rescaling: we state the cluster-level score CLT, the conditions under which the Arellano variance estimator is consistent in the many-fixed-effect regime, and the resulting non-centrality.
\paragraph{Notation.} Partition $\{1,\dots,n\}$ into clusters $\mathcal C_1,\dots,\mathcal C_{G_n}$ with $n_g := |\mathcal C_g|$ and $\bar n_n := \max_g n_g$. Write $\mathcal F^*_n := \sigma(X,\nu,D_{K_n})$, so that conditional on $\mathcal F^*_n$ the residualized regressor $\widetilde X^*_{K_n}=M X^*$ is fixed. Define the cluster loading vectors and their energies
\[
\begin{gathered}
a^{(g)} \in \mathbb R^n, \quad a^{(g)}_i := \widetilde X^{*}_{K_n,i}\,\mathbf 1\{i \in \mathcal C_g\}, \\[3pt]
A_g := \|a^{(g)}\|^2, \qquad \sum_{g} A_g = X^{*\prime}MX^* = \tau^{*2}_n ,
\end{gathered}
\]
the cluster scores $\zeta_g := a^{(g)\prime} u$ with $S_n := \sum_g \zeta_g = X^{*\prime}Mu$, and
\[
\Sigma_g := \operatorname{Var}(u_{\mathcal C_g} \mid \mathcal F^*_n), \qquad
\Psi_n := \operatorname{Var}(S_n \mid \mathcal F^*_n) = \sum_{g} a^{(g)\prime}\Sigma_g\, a^{(g)} .
\]
The Arellano cluster meat \emph{at score scale} and the cluster-robust statistic are
\[
\widehat V^{\mathrm{sc}}_{CR} := \sum_{g}\big(a^{(g)\prime}\widehat u^*\big)^2 ,
\qquad
T^{CR}_n(\beta_0) := \frac{\widehat\beta^*_K - \beta_0}{\sqrt{\widehat V_{CR}}}
= \frac{X^{*\prime}M(Y - X^*\beta_0)}{\sqrt{\widehat V^{\mathrm{sc}}_{CR}}},
\]
where $\widehat V_{CR} = \widehat V^{\mathrm{sc}}_{CR}/\tau^{*4}_n$ is the CRVE of the coefficient and $\widehat\psi := \tau^{*2}_n \widehat V_{CR}/\widehat\sigma^{*2}_{\mathrm{CJN}}$ is the reported variance-inflation factor, i.e.\ the ratio of the squared cluster-robust to the squared i.i.d.\ standard error.
\begin{assumption}[Cluster regularity]\label{ass-cluster}\hfill
\begin{enumerate}[label=(\roman*)]
\item \emph{(Between-cluster independence.)} Conditional on $\mathcal F^*_n$, the subvectors $u_{\mathcal C_1},\dots,u_{\mathcal C_{G_n}}$ are independent with mean zero and arbitrary within-cluster dependence; $\sup_{i,n}\operatorname{\mathbb E}[|u_i|^{2+\delta}\mid\mathcal F^*_n]\le C_u$ for some $\delta \in (0,2]$ and $\sup_{i,n}\operatorname{\mathbb E}[u_i^4\mid\mathcal F^*_n]\le\kappa_u$.
\item \emph{(Many clusters, bounded cluster sizes, no dominant cluster.)} $G_n\to\infty$; the cluster sizes are bounded, $\bar n_n\le\bar n<\infty$ for all $n$; and $\max_g A_g/\tau^{*2}_n\to_p0$. (Given boundedness, the last condition is implied by the maximal-leverage limit in (v), since $\max_gA_g\le\bar n\max_i\widetilde X^{*2}_{K_n,i}$; it is stated separately because it is what the proofs use. Growing cluster sizes are possible under the strengthened conditions of Remark~\ref{rem-clustersize}.)
\item \emph{(Non-degeneracy, and the baseline for $\sigma^2$.)} $n^{-1}\sum_i\operatorname{\mathbb E}[u_i^2\mid\mathcal F^*_n]\to_p\sigma^2\in(0,\infty)$, and $\Psi_n/(\sigma^2\tau^{*2}_n)\to_p\psi\in(0,\infty)$. The first limit is a normalization rather than a restriction, and it is needed: the second alone pins down only the \emph{product} $\sigma^2\psi$, so without a declared baseline the assertion $\psi>1$ has no content. Fixing $\sigma^2$ as the average marginal error variance makes $\psi$ the ratio of the realized score variance to the score variance the same design would carry with independent, homoskedastic errors of that same average variance --- the reading used in Remark~\ref{rem-psi-sign}. Under heteroskedasticity that benchmark is not the independent-error score variance $\sum_i\widetilde X^{*2}_{K_n,i}\operatorname{\mathbb E}[u_i^2\mid\mathcal F^*_n]$, so $\psi$ then mixes dependence with the covariance between the squared residualized regressor and the conditional variance; the theory is unaffected, since only the product $\psi\sigma^2$ enters, but the \emph{interpretation} of $\psi$ as a pure dependence effect requires the homoskedastic reading. Under the conditional homoskedasticity of Theorem~\ref{thm-noncentral} it coincides with the $\sigma^2$ there.
\item \emph{(Projection compatibility.)} $\sum_{g}\big\|Ma^{(g)}-a^{(g)}\big\|^2 = o_p(\tau^{*2}_n)$.
\item \emph{(Design.)} The blocks $\{X_{\mathcal C_g}\}_{g\le G_n}$ are independent across clusters, with unrestricted dependence \emph{within} a cluster; the measurement errors continue to satisfy Assumption~\ref{ass-nu}; and the design-side conclusions of Lemmas~\ref{lem-atten} and~\ref{lem-lev-star} hold:
\[
X'M\nu = o_p(1), \quad \nu'M\nu\to_p c^2(1-\rho), \quad X'MX\to_p(1-\rho)\tau^2, \quad \widehat\lambda_n\to_p\lambda,
\]
\[
\max_i\widetilde\nu^2_{K_n,i}=o_p(1), \qquad \max_i\widetilde X^{*2}_{K_n,i}\big/\tau^{*2}_n\to_p0 ,
\]
the first three giving $X^{*\prime}MX^*\to_p\tau^{*2}=(1-\rho)(\tau^2+c^2)$. The two bias limits are listed separately rather than deduced from the Hessian limits, which would deliver only the combination $2X'M\nu+\nu'M\nu\to_p c^2(1-\rho)$, and it is $X'M\nu+\nu'M\nu$ that the non-centrality needs.
\end{enumerate}
\end{assumption}
Conditions (i)--(iii) are the standard many-clusters apparatus, with cluster sizes bounded. Boundedness is imposed rather than assumed away: the proofs of Lemma~\ref{lem-crve} use it at three separate points, and Remark~\ref{rem-clustersize} states exactly what must be strengthened to allow clusters to grow.
Condition (v) exists because Assumption~\ref{ass-reg}(i) cannot simply be carried over. That condition makes $\{(X_i,u_i)\}$ i.i.d.\ across $i$, which is inconsistent with the within-cluster dependence in $u$ that this section is about, and also excludes the serially dependent \emph{regressor} that drives the empirically common case (Remark~\ref{rem-psi-sign}). What the cluster proofs actually use from the i.i.d.\ theory is not the sampling scheme but the six design-side limits listed in (v), all of which concern $(X,\nu,D_{K_n})$ alone and none of which involves $u$. Lemmas~\ref{lem-atten} and~\ref{lem-lev-star} deliver them under i.i.d.\ sampling. Analogous bounded-cluster primitive conditions can also deliver them with within-cluster dependence in $X$, but bounded cluster size alone is not asserted to be sufficient; the six limits are stated as hypotheses precisely to keep the cluster results free of a stronger sampling claim they do not need.
Condition (iv) is the one with no counterpart in the i.i.d.\ theory, and it is where the many-fixed-effect regime bites. It asks that the fixed-effect projection not distort the cluster loadings, and it fails silently if the fixed effects cut across clusters in a high-dimensional way. The next lemma gives checkable primitive conditions.
\begin{lemma}[When projection compatibility holds]\label{lem-nest}
\begin{enumerate}[label=(\alph*)]
\item \emph{(Nested fixed effects.)} If every fixed-effect cell is contained in a single cluster, then $Ma^{(g)} = a^{(g)}$ \emph{exactly} for every $g$, and Assumption~\ref{ass-cluster}(iv) holds with the left-hand side equal to zero.
\item \emph{(Nested plus a low-dimensional non-nested block.)} Suppose $D_{K_n} = [\,D^{\mathrm{nest}}, D^{\mathrm{ne}}\,]$ with $D^{\mathrm{nest}}$ nested in clusters, and write $\Lambda_n := D^{\mathrm{ne}\prime}M_{\mathrm{nest}}D^{\mathrm{ne}}$. Then
\[
\sum_{g}\big\|Ma^{(g)}-a^{(g)}\big\|^2 \;\le\; \varpi_n \max_g A_g ,
\qquad
\varpi_n := \frac{\operatorname{tr}\big(D^{\mathrm{ne}\prime}D^{\mathrm{ne}}\big)}{\lambda^+_{\min}(\Lambda_n)} ,
\]
with $\lambda^+_{\min}$ the smallest non-zero eigenvalue. The displayed inequality is unconditional. The calibration that follows is not, and needs two balance conditions stated as such: (N1) a \emph{spectral} condition $\lambda^+_{\min}(\Lambda_n)\asymp n/d^{\mathrm{ne}}_n$ with $d^{\mathrm{ne}}_n:=\operatorname{rank}(D^{\mathrm{ne}})$, which together with $\operatorname{tr}(D^{\mathrm{ne}\prime}D^{\mathrm{ne}})\asymp n$ gives $\varpi_n\asymp d^{\mathrm{ne}}_n$; and (N2) an \emph{energy} condition $\max_gA_g = O_p(\tau^{*2}_n/G_n)$. Under (N1)--(N2), \emph{condition (iv) reduces to $d^{\mathrm{ne}}_n/G_n \to 0$}: the number of fixed effects that cut across clusters must be small relative to the number of clusters. Neither condition follows from cell counts alone. (N1) holds exactly in the balanced two-way panel --- with $G$ units and $T$ periods, $\Lambda_n = G(I_T-T^{-1}\mathbf 1\mathbf 1')$, so $\lambda^+_{\min}=G=n/T$ --- but equal cell \emph{sizes} do not by themselves control the smallest non-zero eigenvalue of the residualized Gram matrix. (N2) concerns the residualized regressor rather than the design: equal cluster sizes with $\widetilde X^*$ energy concentrated in $\sqrt{G_n}$ clusters give $\max_gA_g/\tau^{*2}_n\asymp G_n^{-1/2}$, and the reduction then requires $d^{\mathrm{ne}}_n/\sqrt{G_n}\to0$ instead. Both are computable from the design and the residualized regressor before any inference, which is the point of stating them.
\end{enumerate}
\end{lemma}
\begin{proof}
Deferred to \ref{app-cluster}.
\end{proof}
Part (b) is the practically binding condition and it is computable before any estimation. In the lead application (Section~\ref{sec-app-vdem}) the country effects are nested in country clusters while the $59$ year effects are not, against $G_n=163$ clusters, so $d^{\mathrm{ne}}_n/G_n\approx0.36$; in the second application (Section~\ref{sec-app-psid}) the person effects are nested in person clusters and the $7$ year effects are not, against $G_n=595$, giving $d^{\mathrm{ne}}_n/G_n\approx0.012$. The cluster-robust columns of the two applications therefore do not carry the same warrant, and we say so where they are reported.
\begin{lemma}[Cluster score CLT]\label{lem-cluster-clt}
Under Assumption~\ref{ass-nu}, Assumption~\ref{ass-cluster}(i)--(iii) and (v), and the drift~\eqref{eq-local-drift},
\[
\frac{X^{*\prime}M u}{\sqrt{\Psi_n}} \;\Rightarrow\; N(0,1).
\]
\end{lemma}
\begin{lemma}[Consistency of the cluster-robust variance estimator]\label{lem-crve}
Under the conditions of Lemma~\ref{lem-cluster-clt} and Assumption~\ref{ass-cluster}(iv),
\[
\widehat V^{\mathrm{sc}}_{CR}\big/\Psi_n \;\to_p\; 1 .
\]
The statement is for the \emph{uncorrected} meat. Applying the conventional small-sample factor $\tfrac{G_n}{G_n-1}\cdot\tfrac{n-1}{n-K_n}$, $K_n = d_{K_n}+1$, multiplies the estimator by a factor converging to $1/(1-\rho)$ rather than to one, and so \emph{over}-corrects whenever $\rho$ is non-negligible. Under nesting there is no residual-shrinkage bias to correct --- part (a) of Lemma~\ref{lem-nest} makes $a^{(g)\prime}\widehat u^*$ an exact linear functional of the errors with no projection loss --- and the $\tfrac{n-1}{n-K_n}$ factor should be omitted.
\end{lemma}
\begin{remark}[Growing cluster sizes]\label{rem-clustersize}
Assumption~\ref{ass-cluster}(ii) imposes bounded cluster sizes, and the proofs use boundedness at three identifiable points, all in Lemma~\ref{lem-crve}: the conditional-covariance operator-norm bound $\|\operatorname{Var}(u\mid\mathcal F^*_n)\|_{\mathrm{op}}\le\bar n_n\kappa_u^{1/2}$ in Step 2(2a); the bound $\max_g\sum_{i\in\mathcal C_g}\widetilde\nu^2_{K_n,i}\le\bar n_n\max_i\widetilde\nu^2_{K_n,i}$ in Step 2(2b); and the fourth-moment bound $\operatorname{\mathbb E}[\zeta_g^4\mid\mathcal F^*_n]\le\kappa_u\bar n_n^2A_g^2$ in Step 1. Clusters may be allowed to grow, $\bar n_n\to\infty$, provided each of these is restored explicitly:
\begin{enumerate}[label=(\roman*$'$)]
\item \emph{(CLT.)} $\bar n_n^{(2+\delta)/\delta}\,\max_gA_g/\tau^{*2}_n\to_p0$, which is what the Lyapunov bound of Lemma~\ref{lem-cluster-clt} actually requires;
\item \emph{(Step 1.)} $\bar n_n^{2}\,\max_gA_g/\tau^{*2}_n\to_p0$, implied by (i$'$) when $\delta\le2$;
\item \emph{(Step 2a.)} $\bar n_n\sum_g\|Ma^{(g)}-a^{(g)}\|^2 = o_p(\tau^{*2}_n)$, a strengthening of Assumption~\ref{ass-cluster}(iv) by the factor $\bar n_n$;
\item \emph{(Step 2b.)} $\bar n_n\max_i\widetilde\nu^2_{K_n,i}=o_p(1)$, which by Lemma~\ref{lem-lev-star} holds whenever $\bar n_n = o(n^{1-1/r})$.
\end{enumerate}
We state the bounded-size version in the text because it is the empirically relevant one --- both applications have $\bar n_n\le59$ against $G_n$ of $163$ and $595$ --- and because the projection-compatibility condition (iv) alone does \emph{not} absorb the extra $\bar n_n$ in (iii$'$), so advertising ``bounded or slowly growing'' without (i$'$)--(iv$'$) would leave a gap between the statement and the proof.
\end{remark}
\begin{theorem}[Cluster-robust non-centrality]\label{thm-cluster}
Under Assumption~\ref{ass-nu} with $r=2$, Assumption~\ref{ass-cluster}, the local drift~\eqref{eq-local-drift}, $0<\tau^2<\infty$, and $H_0:\beta=\beta_0$,
\[
T^{CR}_n(\beta_0) \;\Rightarrow\; N\big(\eta_{CR},\,1\big),
\qquad
\eta_{CR} \;=\; \frac{\eta}{\sqrt\psi} \;=\; -\frac{\beta_0\,c^2\sqrt{1-\rho}}{\sigma\sqrt{\psi}\,\sqrt{\tau^2+c^2}} ,
\]
with $\eta$ as in Theorem~\ref{thm-noncentral}. Cluster dependence therefore enters the diagnostic through the denominator alone: the attenuation bias in the numerator is $\operatorname{\mathbb E}[\nu'M_{K_n}\nu]\beta_0$, a functional of the measurement error only, and $\nu\perp u$ by Assumption~\ref{ass-nu}(i), so no dependence structure in $u$ can touch it.
\end{theorem}
Three objects must now be kept apart, because they have different statuses and the distinction is exactly where a careless argument goes wrong. The \emph{oracle} non-centrality is evaluated at the true $\beta_0$; the \emph{point plug-in} diagnostic is evaluated at the corrected pilot $\widehat\beta_0^{\mathrm{corr}}$, which is \emph{not} consistent in the weak-information regime (Corollary~\ref{cor-slope}, Remark~\ref{rem-corr-noise}); and the \emph{formal certificate} is evaluated at an upper confidence bound for $|\beta_0|$. We take them in that order.
\begin{corollary}[Oracle rescaling, and the algebraic cancellation]\label{cor-cluster-feasible}
Define the \emph{cluster-robust scale}
\[
\widehat s^2_{CR} \;:=\; \widehat\sigma^{*2}_{\mathrm{CJN}}\,\widehat\psi \;=\; \widehat V^{\mathrm{sc}}_{CR}\big/\tau^{*2}_n ,
\]
the second equality being an algebraic identity, not an approximation: $\widehat\sigma^{*2}_{\mathrm{CJN}}$ cancels exactly. Under the conditions of Theorem~\ref{thm-cluster}:
\begin{enumerate}[label=(\alph*)]
\item $\widehat s^2_{CR}\to_p\psi\sigma^2$. This holds even though $\widehat\sigma^{*2}_{\mathrm{CJN}}$ is \emph{not} consistent for $\sigma^2$ under cluster dependence, and $\widehat\psi$ is \emph{not} consistent for $\psi$: \emph{if} $\widehat\sigma^{*2}_{\mathrm{CJN}}\to_p\varsigma^2\in(0,\infty)$ --- an extra requirement, since nothing in Assumption~\ref{ass-cluster} forces this estimator to have a limit --- then $\widehat\psi\to_p\psi\sigma^2/\varsigma^2$, and the two errors are reciprocal. Part~(a) itself does not need $\varsigma^2$ to exist: $\widehat s^2_{CR}=\widehat V^{\mathrm{sc}}_{CR}/\tau^{*2}_n$ never mentions $\widehat\sigma^{*2}_{\mathrm{CJN}}$.
\item \emph{(Oracle.)} At the true $\beta_0$, the feasible non-centrality formed with $\widehat s_{CR}$ in place of $\widehat\sigma^*_{\mathrm{CJN}}$ is consistent for the cluster-robust target:
\[
\frac{|\beta_0|}{\widehat s_{CR}}\,(1-\widehat\lambda_n)\sqrt{\tau^{*2}_n} \;\to_p\; \frac{|\beta_0|}{\sigma\sqrt\psi}(1-\lambda)\sqrt{\tau^{*2}} \;=\; \frac{|\eta|}{\sqrt\psi} \;=\; |\eta_{CR}| .
\]
More generally, for any deterministic $b$ substituted for $|\beta_0|$, the same statement holds with $|\eta|$ replaced by its value at $b$.
\item \emph{(Breakdown, exactly.)} $t^*_n/\sqrt{\widehat\psi}$ is \emph{identically} the reported cluster-robust $t$-statistic $t^{CR}_n := |\widehat\beta^*_K|/\sqrt{\widehat V_{CR}}$, so
\[
\lambda^{\dagger}_{CR} \;=\; \frac{t^{CR}_n}{\,t^{CR}_n+\eta^{\dagger}(\alpha,\delta)\,}
\]
is Definition~\ref{def-breakdown} evaluated at the cluster-robust $t$. This is an algebraic identity requiring no limit theory, and it is the operational content of the rescaling: \emph{the cluster-robust diagnostic is the i.i.d.\ diagnostic run on the reported cluster-robust $t$-statistic.}
\end{enumerate}
\end{corollary}
\begin{remark}[The point plug-in is descriptive, not consistent]\label{rem-plugin}
Substituting the corrected pilot $\widehat\beta_0^{\mathrm{corr}}=\widehat\beta^*_K/\widehat\lambda_n$ for $|\beta_0|$ in Corollary~\ref{cor-cluster-feasible}(b) does \emph{not} yield a consistent estimator of $|\eta_{CR}|$, and no argument in this paper claims otherwise. In the weak-information regime $\tau^2<\infty$ the pilot is asymptotically centred at $\beta_0$ but nondegenerate. Its limiting variance is \emph{not} the i.i.d.\ one of Corollary~\ref{cor-slope} and Proposition~\ref{prop-pilot}(ii): under cluster dependence the score variance is $\Psi_n\to_p\psi\sigma^2\tau^{*2}$ rather than $\sigma^2\tau^{*2}$, so the same Slutsky argument applied to Lemma~\ref{lem-cluster-clt} in place of the i.i.d.\ score CLT gives
\[
\widehat\beta_0^{\mathrm{corr}}\;\Rightarrow\;B \;:=\; \beta_0+N\big(0,\ \psi\,\sigma^2/(\lambda^2\tau^{*2})\big) ,
\]
the extra factor $\psi$ being exactly what makes $\widehat{\mathrm{se}}(\widehat\beta_0^{\mathrm{corr}})=\widehat s_{CR}/(\widehat\lambda_n\sqrt{\tau^{*2}_n})$ the right standard error for it in Proposition~\ref{prop-certificate}. Consequently, by the continuous mapping theorem and part (a), when $\beta_0\ne0$,
\[
|\widehat\eta_{CR}| \;:=\; \frac{|\widehat\beta_0^{\mathrm{corr}}|}{\widehat s_{CR}}(1-\widehat\lambda_n)\sqrt{\tau^{*2}_n}
\;\Rightarrow\; \frac{|B|}{|\beta_0|}\,|\eta_{CR}| ,
\]
a nondegenerate random multiple of the target, with median close to it but no concentration. (At $\beta_0=0$ the ratio is undefined and the statement is instead that $|\eta_{CR}|=0$ while $|\widehat\eta_{CR}|\Rightarrow|B|(1-\lambda)\sqrt{\tau^{*2}}/(\sigma\sqrt\psi)$ with $B$ centred at zero: the plug-in remains nondegenerate even though the target vanishes, which is the same point.) This is precisely why the plug-in verdict is labelled a \emph{point pass} and not a certificate: it is a descriptive statistic, and its sampling variability is a first-order feature of the weak-information regime rather than an approximation error that vanishes. Consistency is recovered only under strong information $nQ_{K_n}\to\infty$, where the Gaussian term disappears. A size-controlled statement requires Proposition~\ref{prop-certificate}.
\end{remark}
\begin{proposition}[Formal certificate]\label{prop-certificate}
Fix $0<\alpha<1$, a size tolerance $0<\delta<1-\alpha$, and a separate confidence level $\gamma\in(0,\tfrac12]$ (so that $z_{1-\gamma}\ge0$ and the bound $U_n$ below is an inflation of the pilot rather than a deflation; a certificate at $\gamma>\tfrac12$ would not be a conservative object in any case). On samples with $0<\widehat\lambda_n\le1$, let
\[
U_n \;:=\; \big|\widehat\beta_0^{\mathrm{corr}}\big| \;+\; z_{1-\gamma}\,\widehat{\mathrm{se}}\big(\widehat\beta_0^{\mathrm{corr}}\big),
\qquad
\widehat{\mathrm{se}}\big(\widehat\beta_0^{\mathrm{corr}}\big) \;=\; \frac{\widehat s_{CR}}{\widehat\lambda_n\sqrt{\tau^{*2}_n}} ,
\]
and define the \emph{certified non-centrality bound} and the certification rule
\[
\widehat\eta^{\,U}_n \;:=\; \frac{U_n}{\widehat s_{CR}}\,(1-\widehat\lambda_n)\sqrt{\tau^{*2}_n} ,
\qquad
\text{\emph{certify} } \iff \widehat\eta^{\,U}_n \le \eta^{\dagger}(\alpha,\delta) .
\]
If $\widehat\lambda_n\le0$, the rule does not certify; a non-positive estimated reliability is itself an automatic flag. Since $\widehat\lambda_n\to_p\lambda>0$, this convention affects an event of probability tending to zero. A non-negative variance pilot already ensures $\widehat\lambda_n\le1$; otherwise truncate it at one.
Under the conditions of Theorem~\ref{thm-cluster}:
\begin{enumerate}[label=(\alph*)]
\item $U_n$ is an asymptotically valid one-sided $(1-\gamma)$ upper confidence bound for $|\beta_0|$: $\liminf_n\mathbb P\big(|\beta_0|\le U_n\big)\ge1-\gamma$.
\item $\widehat\eta^{\,U}_n$ is an asymptotically valid $(1-\gamma)$ upper confidence bound for $|\eta_{CR}|$.
\item Consequently the certification rule has asymptotic false-certification probability at most $\gamma$:
\[
\limsup_n\ \mathbb P\Big(\text{certify, yet the true asymptotic size exceeds } \alpha+\delta\Big) \;\le\; \gamma .
\]
\end{enumerate}
The i.i.d.\ version is the special case $\widehat\psi\equiv1$, $\widehat s_{CR}=\widehat\sigma^*_{\mathrm{CJN}}$.
\end{proposition}
\begin{remark}[How $\gamma$ and $\delta$ combine, and how they do not]\label{rem-gamma-delta}
The two are separate error budgets answering different questions, and Proposition~\ref{prop-certificate} deliberately does not merge them. The tolerance $\delta$ bounds the object being certified --- the size distortion of the conventional $t$-test. The level $\gamma$ bounds the probability that the certification \emph{statement} is wrong, i.e.\ that a specification is certified whose true distortion exceeds $\delta$. They should be reported as a triple $(\alpha,\delta,\gamma)$ and not collapsed.
A single unconditional number can be extracted if one insists, but it is weaker and less informative. Writing $S$ for the true asymptotic size of the nominal-$\alpha$ test, Proposition~\ref{prop-certificate}(c) gives $\mathbb P(\text{certify},\,S>\alpha+\delta)\le\gamma$; since $S\le1$ always, the unconditional bound on the size of a certified specification is
\[
\operatorname{\mathbb E}\big[S\,\big|\,\text{certify}\big] \;\le\; \alpha+\delta+\gamma\big(1-\alpha-\delta\big)\Big/\mathbb P(\text{certify}) ,
\]
which degrades quickly when certification is rare. Thus the theorem controls the \emph{unconditional} probability of issuing a false certificate; it does not assert that the conditional false-certificate probability given certification is at most $\gamma$ (the latter is bounded only by $\gamma/\mathbb P(\text{certify})$). Throughout the applications we take $\gamma = 0.05$ alongside $\delta = 0.05$; nothing forces the two to be equal, and a reader who regards pilot uncertainty as the greater risk should lower $\gamma$ rather than $\delta$.
\end{remark}
\begin{proof}
Lemmas~\ref{lem-cluster-clt}--\ref{lem-crve}, Theorem~\ref{thm-cluster}, Corollary~\ref{cor-cluster-feasible} and Proposition~\ref{prop-certificate} are proved in \ref{app-cluster}.
\end{proof}
\begin{remark}[The sign of $\psi-1$ is not free, and not what one might guess]\label{rem-psi-sign}
It is tempting to reason that serial dependence in $u$ inflates standard errors, hence $\psi>1$, hence clustering mechanically deflates the measured distortion and the i.i.d.\ verdict is conservative. That reasoning is incomplete, and its conclusion can fail. Because $\Psi_n = \sum_g a^{(g)\prime}\Sigma_g a^{(g)}$ weights the within-cluster covariance by the \emph{residualized} regressor, and because $\widetilde X^{*}_{K_n}$ is orthogonal to every fixed-effect dummy, an equicorrelated component of $u$ at a level that is itself a fixed effect is annihilated exactly, as in Proposition~\ref{prop-equicorr}. With a serially dependent error but a within-cluster \emph{serially independent} regressor, the surviving terms are predominantly negative and $\psi<1$: clustering then makes the diagnostic \emph{more} alarming, not less. What drives $\psi$ above one is persistence in the regressor and the error \emph{together}, over a panel long enough for that persistence to accumulate. Simulations on a unit-and-time FE design confirm this: at $G=200$, an i.i.d.\ treatment with AR(1) errors ($\varrho_u=0.6$, $T=10$) gives $\psi\approx0.75$, while $\varrho_x=0.9,\varrho_u=0.9$ at $T=20$ gives $\psi\approx1.27$ and $\varrho_x=0.98,\varrho_u=0.95$ at $T=40$ gives $\psi\approx1.72$ (Section~\ref{sec-design5}(vi); the panel length rises with the persistence across the three, and $\psi$ depends on both).
Two cautions follow. First, the empirically common case --- a persistent regressor such as a democracy index or a log wage, together with persistent errors --- is the $\psi>1$ case, but that is an empirical fact about those designs and not a theorem. Second, and more easily missed, $\psi$ and its reported estimate $\widehat\psi$ are different objects whose comparisons with one need not agree: by Corollary~\ref{cor-cluster-feasible}(a), $\widehat\psi\to_p\psi\sigma^2/\varsigma^2$, and $\varsigma^2<\sigma^2$ whenever the fixed-effect projection removes positively correlated error. The gap is not small. In the nested configuration of Section~\ref{sec-design5} the population $\psi$ is $0.64$ --- clustering genuinely amplifies the distortion there --- while the reported $\widehat\psi$ sits at $0.99$. So $\psi<1$ does not show up as $\widehat\psi<1$, and one may not read the population statement off the reported number or the reverse. What settles the applied verdict is $\widehat\psi$, not $\psi$: by Corollary~\ref{cor-cluster-feasible}(c) the cluster-robust diagnostic is an exact function of $\widehat\psi$ in the sample, whereas $\psi$ enters only the population target. The direction of the adjustment must therefore be read off $\widehat\psi$, and this paragraph's mechanism read as an account of $\psi$.
\end{remark}
\section{Critical Values}\label{sec-cv}
\subsection{Threshold and feasible form}
We invert the leading (quadratic) size distortion to obtain a Stock--Yogo-style threshold.
\begin{corollary}[Critical-value formula]\label{cor-cv}
Under the conditions of Theorem~\ref{thm-noncentral}, the asymptotic size of the nominal-$\alpha$ test $|T_n^{\mathrm{CJN}*}| > z_{1-\alpha/2}$ is \emph{even} in $\eta$ with a vanishing first derivative, so its leading distortion is \emph{quadratic}:
\[
\mathbb P(|T_n^{\mathrm{CJN}*}| > z_{1-\alpha/2}) = \alpha + z_{1-\alpha/2}\,\phi(z_{1-\alpha/2})\,\eta^2 + O(\eta^4)
\]
as $|\eta| \to 0$. The critical value bounding the leading distortion below $\delta$ is, writing $z := z_{1-\alpha/2}$,
\begin{equation}\label{eq-cv}
\tau^2_{\mathrm{crit}}(\rho, c^2, \beta_0, \sigma; \alpha, \delta)
= \frac{\beta_0^2\, c^4\,(1-\rho)\, z\,\phi(z)}{\sigma^2\delta} - c^2,
\end{equation}
positive whenever $\beta_0^2\, c^2 (1-\rho)\, z\,\phi(z) > \sigma^2\delta$. Two features of the formula trace directly to the Hessian limit of Lemma~\ref{lem-clt} and are worth flagging because they are easy to get wrong: the leading term is linear in $(1-\rho)$, not quadratic, and the subtracted term is $c^2$, not $c^2(1-\rho)$ (Remark~\ref{rem-rho}).
\end{corollary}
\begin{proof}[Proof sketch]
For $T\sim N(\eta,1)$ the two-sided rejection probability $s(\eta)=\Phi(-z-\eta)+1-\Phi(z-\eta)$ is even with $s'(0)=0$ and $s''(0)=2z\phi(z)$, so $s(\eta)=\alpha+z\phi(z)\eta^2+O(\eta^4)$; substituting $\eta^2=\beta_0^2 c^4(1-\rho)/(\sigma^2(\tau^2+c^2))$ (Theorem~\ref{thm-noncentral}) and setting the leading term equal to $\delta$ rearranges to~\eqref{eq-cv}. Full derivation in \ref{app-cv}.
\end{proof}
\begin{remark}[Quadratic versus exact-inversion threshold]\label{rem-exact-cv}
Formula~\eqref{eq-cv} inverts the leading \emph{quadratic} distortion $z\phi(z)\eta^2$, equivalently the threshold $|\eta| \le \eta^{\mathrm{quad}} := \sqrt{\delta/(z\phi(z))}$. For $0<\alpha<1$ and $0<\delta<1-\alpha$, the \emph{exact-inversion} threshold replaces $\eta^{\mathrm{quad}}$ by $\eta^{\dagger}(\alpha,\delta)$, the unique positive root of $\Phi(-z-\eta)+1-\Phi(z-\eta)=\alpha+\delta$, giving
\begin{equation}\label{eq-cv-exact}
\tau^{2,\mathrm{exact}}_{\mathrm{crit}} = \left(\frac{|\beta_0| c^2\sqrt{1-\rho}}{\sigma\,\eta^{\dagger}(\alpha,\delta)}\right)^2 - c^2.
\end{equation}
For either formula, the operational threshold is the positive part of the displayed algebraic boundary: if the right-hand side is negative, every admissible $\tau^2\ge0$ passes that inversion.
At $\alpha=\delta=0.05$, $\eta^{\mathrm{quad}}=0.661$ versus $\eta^{\dagger}=0.652$: the quadratic closed form is within $1.4\%$ of the exact root in $\eta$. Because $\eta^{\mathrm{quad}}>\eta^{\dagger}$ it is very slightly \emph{anti-conservative} (at its $\delta=0.05$ boundary the true distortion is $0.0513$, just above the nominal $\delta$), so it is an accurate approximation rather than a strict upper bound. This is a different and far milder failure than the discarded linear surrogate $\delta/(2\phi)=0.428$, which over-demanded $\tau^2$ by a factor $(\eta^{\dagger}/0.428)^2\approx 2.3$ and, worse, \emph{fails to control size} at larger $\delta$ (at $\delta=0.15$ its boundary $\eta=1.28$ carries a true distortion of $0.20$). The quadratic form~\eqref{eq-cv} degrades further only in the far tail $\eta \gtrsim 1$ ($\lambda \lesssim 0.5$), where the exact inversion~\eqref{eq-cv-exact} should be used; Design~2 of Section~\ref{sec-sim} confirms that the two nearly coincide across $(n,\rho,c^2)$ at the moderate $\delta$ the paper targets. We recommend~\eqref{eq-cv-exact} as headline, with~\eqref{eq-cv} as its closed-form companion.
\end{remark}
The threshold has an equivalent form that is more useful in practice, because it removes the drift-scale quantities $c^2$ and $\tau^2$ in favor of directly observable objects.
\begin{corollary}[Feasible non-centrality]\label{cor-feasible}
Under the hypotheses of Lemma~\ref{lem-atten}, let $\tau^{*2}_n := X^{*\prime} M_{K_n} X^*$ denote the \emph{observed} (random) residual variation of the contaminated regressor --- distinct from its deterministic limit $\tau^{*2} = (1-\rho)(\tau^2 + c^2)$, to which it converges, $\tau^{*2}_n \to_p \tau^{*2}$ --- and $\lambda_n$ the within reliability of~\eqref{eq-lambda}. \emph{This corollary and everything built on it (Definition~\ref{def-breakdown}, the protocol of Section~\ref{sec-protocol}, and the empirical applications of Section~\ref{sec-app}) never requires the practitioner to decompose the diagnostic into $(\rho,\tau^2,c^2)$: $\tau^{*2}_n$ is observed from the regression, while $\lambda_n$ is supplied by the external reliability or noise pilot. The feasible diagnostic is therefore insulated from the primitive decomposition once that pilot is specified.} The population non-centrality obeys the exact identity $|\eta| = (|\beta_0|/\sigma)(1-\lambda)\sqrt{\tau^{*2}}$, and the plug-in
\[
|\widehat\eta_n| := \frac{|\beta_0|}{\sigma}\,(1 - \lambda_n)\, \sqrt{\tau^{*2}_n} \quad\text{satisfies}\quad |\widehat\eta_n| = |\eta|\,(1 + o_p(1)),
\]
so the size condition ``leading distortion $\le \delta$'' is asymptotically equivalent to
\begin{equation}\label{eq-feasible}
\frac{|\beta_0|}{\sigma}\,(1 - \lambda_n)\, \sqrt{\tau^{*2}_n} \;\le\; \sqrt{\frac{\delta}{z_{1-\alpha/2}\,\phi(z_{1-\alpha/2})}}\;=\;\eta^{\mathrm{quad}},
\end{equation}
its exact-normal-inversion analogue being asymptotically equivalent to $(|\beta_0|/\sigma)(1-\lambda_n)\sqrt{\tau^{*2}_n} \le \eta^{\dagger}(\alpha,\delta)$ (both feasible forms carry a $1+o_p(1)$ from replacing $\tau^{*2}$ by the random $\tau^{*2}_n$).
\end{corollary}
\begin{proof}[Proof sketch]
The identity $|\eta| = (|\beta_0|/\sigma)(1-\lambda)\sqrt{\tau^{*2}}$ is algebraic: with $\lambda=\tau^2/(\tau^2+c^2)$ and $\tau^{*2}=(1-\rho)(\tau^2+c^2)$, $(1-\lambda)\sqrt{\tau^{*2}} = \frac{c^2}{\tau^2+c^2}\sqrt{(1-\rho)(\tau^2+c^2)} = \frac{c^2\sqrt{1-\rho}}{\sqrt{\tau^2+c^2}}$, which matches $\eta$ of Theorem~\ref{thm-noncentral} up to the sign and the $|\beta_0|/\sigma$ factor. Since $X^{*\prime}M_{K_n}X^*\to_p\tau^{*2}$ by Lemma~\ref{lem-atten}, rearranging $z\phi(z)\eta^2\le\delta$ of Corollary~\ref{cor-cv} gives~\eqref{eq-feasible}. See \ref{app-feasible}.
\end{proof}
Equation~\eqref{eq-feasible} is the form we recommend reporting: $\tau^{*2}_n$ is printed by any FE regression, $\sigma$ is the CJN residual standard deviation, and $(1-\lambda_n)$, the within noise share, is the single quantity requiring external input. At $\alpha = 0.05$, $\eta^{\mathrm{quad}}=\sqrt{\delta/(z\phi(z))}=\sqrt{8.73\,\delta}$: for instance, tolerating one percentage point of size distortion ($\delta = 0.01$) requires $(|\beta_0|/\sigma)(1-\lambda_n)\sqrt{\tau^{*2}_n} \le 0.296$ (the exact-inversion bound $\eta^{\dagger}(0.05,0.01)=0.295$ nearly coincides, confirming the quadratic form).
\subsection{Pilots and the correction}\label{sec-pilot}
The threshold depends on four quantities. Only $\rho_n = d_{K_n}/n$ and $\tau^{*2}_n$ are directly observable, and $\sigma$ is estimated by $\widehat\sigma^*_{\mathrm{CJN}}$ (consistent by Theorem~\ref{thm-noncentral}, Step 3). The remaining inputs are $|\beta_0|$ and the noise level. Two warnings are in order, the first of which is substantive.
\begin{proposition}[The naive pilot is anti-conservative; the corrected pilot]\label{prop-pilot}
Under the conditions of Theorem~\ref{thm-noncentral}, suppose the noise pilot $\widehat{a}_n \to_p c^2(1-\rho)$ (equivalently $\widehat\lambda_n := \tau^{*2}_n{}^{-1}(\tau^{*2}_n - \widehat a_n) \to_p \lambda$), obtained from any of the sources in Remark~\ref{rem-est}. Then:
\begin{enumerate}[label=(\roman*)]
\item \emph{(Anti-conservative centering.)} By Corollary~\ref{cor-slope} the naive point estimate is asymptotically centered at the attenuated value $\beta_0\lambda$ --- the center of its weak-information limiting law, where $\lambda < 1$ and attenuation bites; under strong information $\lambda \to 1$ and the attenuation vanishes --- so $|\widehat\beta^*_K|$ is centered at $\lambda|\beta_0| < |\beta_0|$. Since the distortion $z\phi(z)\eta^2$ is strictly increasing in $|\beta_0|$, evaluating it at this center understates the true distortion by the factor $\lambda^2$, so the naive diagnostic is biased toward \emph{passing} specifications that fail the true threshold. Two caveats bound this statement to the \emph{center}: under weak information $|\widehat\beta^*_K|$ is a folded normal, not a constant, so its sampling spread is a separate source of uncertainty (addressed below); and because the limiting law of $\widehat\beta^*_K$ has second moment $\lambda^2\beta_0^2+\sigma^2/\tau^{*2}$, a \emph{squared} plug-in must use the recentered $\widehat\beta_0^{\mathrm{corr}}$ rather than $\widehat\beta^{*2}_K$, whose limit is inflated by the variance term.
\item \emph{(Bias-corrected pilot.)} The reliability-corrected pilot
\[
\widehat\beta_0^{\mathrm{corr}} := \frac{\widehat\beta^*_K}{\widehat\lambda_n}
\]
removes the attenuation bias: $\widehat\beta_0^{\mathrm{corr}} \Rightarrow \beta_0 + N(0,\sigma^2/(\lambda^2\tau^{*2}))$, i.e.\ it is asymptotically centered at $\beta_0$. In the strong-information regime $nQ_{K_n}\to\infty$ the Gaussian term vanishes and $\widehat\beta_0^{\mathrm{corr}}\to_p\beta_0$ (consistent). In the weak-information regime $\tau^2<\infty$ it is centered at $\beta_0$ but carries sampling error of order $\sigma/(\lambda\sqrt{\tau^{*2}})$; this is a coefficient-uncertainty source \emph{distinct} from reliability uncertainty, and controlling it requires a confidence interval (or worst-case bound) for $\beta_0$, not merely the reliability band of Remark~\ref{rem-corr-noise}. Both vanish under strong information.
\end{enumerate}
\end{proposition}
\begin{proof}[Proof sketch]
(i) follows from the centering of $\widehat\beta^*_K$ at $\beta_0\lambda$ (Corollary~\ref{cor-slope}) and the strict monotonicity of $z\phi(z)\eta^2$ in $|\beta_0|$. (ii) With $\widehat\beta^*_K\Rightarrow\beta_0\lambda+N(0,\sigma^2/\tau^{*2})$ and $\widehat\lambda_n\to_p\lambda>0$, Slutsky gives $\widehat\beta_0^{\mathrm{corr}}\Rightarrow\beta_0+N(0,\sigma^2/(\lambda^2\tau^{*2}))$, a constant $\beta_0$ under strong information. See \ref{app-pilot}.
\end{proof}
\begin{remark}[Sampling noise of the corrected pilot at small $\lambda$]\label{rem-corr-noise}
The correction $\widehat\beta_0^{\mathrm{corr}}=\widehat\beta^*_K/\widehat\lambda_n$ divides by an estimated reliability, so its sampling variance is inflated by roughly $\widehat\lambda_n^{-2}$ (delta method), and the inflation is largest precisely in the low-reliability regime where the diagnostic bites. Two distinct uncertainties must therefore be separated. \emph{Reliability} uncertainty (how noisy is the regressor?) is handled by reporting the diagnostic over an interval of $\lambda$ values, the breakdown-reliability band of Remark~\ref{rem-est}(iii) and Definition~\ref{def-breakdown}; the empirical section does this. \emph{Coefficient} uncertainty is separate and is \emph{not} controlled by the $\lambda$-band: by Corollary~\ref{cor-slope}, $\widehat\beta_0^{\mathrm{corr}} \Rightarrow \beta_0 + N(0,\sigma^2/(\lambda^2\tau^{*2}))$ is nondegenerate under weak information, so a point plug-in of $|\widehat\beta_0^{\mathrm{corr}}|$ yields a \emph{descriptive} verdict, not a size-controlled certificate. For a formally conservative certificate one evaluates the non-centrality at an \emph{upper} confidence bound $|\widehat\beta_0^{\mathrm{corr}}| + z_{1-\gamma}\,\widehat{\mathrm{se}}(\widehat\beta_0^{\mathrm{corr}})$ in place of the point estimate, at a confidence level $\gamma$ named separately from the size tolerance $\delta$; Proposition~\ref{prop-certificate} shows this controls the false-certification probability at $\gamma$, and Remark~\ref{rem-gamma-delta} explains why $\gamma$ and $\delta$ are not combined. Both uncertainties vanish under strong information $nQ_{K_n}\to\infty$. In the applications we report the point verdicts and note that the flagged specifications fail by margins far exceeding this coefficient uncertainty.
\end{remark}
Proposition~\ref{prop-pilot}(i) is the point made in Section~\ref{sec-notjust}: the most natural implementation of the diagnostic is biased in the dangerous direction, by exactly the reliability $\lambda$ whose smallness motivates running it. The correction costs one line.
\begin{remark}[Sources for the noise pilot]\label{rem-est}
For $\sigma_\nu^2$ (and hence $c^2=n\sigma_\nu^2$ at the sample's local-drift scale, or the equivalent within noise share $1-\lambda_n$), three sources are available in decreasing order of directness. \emph{(i) Measurement-model posteriors}: producers of some expert-coded or latent-trait indices publish observation-level posterior uncertainty, as the V-Dem model and the Unified Democracy Scores of \citet{PMM} do, making $\widehat\sigma_\nu^2$ directly available; this is the basis of the lead application. \emph{(ii) Validated survey or administrative-matched measures}: with independent repeated measures $X^*_{i,1},X^*_{i,2}$ of the same $X_i$, $\widehat\sigma_\nu^2=\tfrac12\operatorname{\mathbb E}_i[(X^*_{i,1}-X^*_{i,2})^2]$. Reinterview and test--retest studies supply such repeats; a survey measure matched to an administrative benchmark supplies a validation discrepancy from which reliability can be estimated, subject to any error in the benchmark. Examples include the CPS--Social Security earnings match of \citet{BK} and the PSID Validation Study of \citet{BoundEtAl94}, the latter reporting within- and difference-transformed reliabilities. \emph{(iii) Sensitivity bands}: absent external information, evaluate the diagnostic over a grid of $\lambda$ and report the breakdown value (Definition~\ref{def-breakdown}). Thus the point diagnostic is usable beyond specialist expert-coded data whenever a defensible reliability range exists; a formal certificate still requires a defensible lower reliability bound. A sensitivity grid alone describes where the verdict changes but cannot manufacture that external information.
\end{remark}
\begin{definition}[Breakdown reliability]\label{def-breakdown}
Fix $0<\alpha<1$ and $0<\delta<1-\alpha$. The \emph{breakdown reliability} $\lambda^{\dagger}$ of a specification is the minimum reliability at which the feasible size condition~\eqref{eq-feasible}, evaluated with the reliability-corrected pilot $\widehat\beta_0^{\mathrm{corr}} = \widehat\beta^*_K/\lambda$, holds --- i.e.\ the fixed point $\lambda = \lambda^{\dagger}(\lambda)$ of the required-reliability map. Writing $t^*_n := |\widehat\beta^*_K|\sqrt{\tau^{*2}_n}/\widehat\sigma^*_{\mathrm{CJN}}$ (the specification's conventional $t$-statistic), the fixed point has the closed form
\[
\lambda^{\dagger} := \frac{t^*_n}{\,t^*_n + \eta^{\dagger}(\alpha,\delta)\,},
\]
with the quadratic version replacing $\eta^{\dagger}$ by $\sqrt{\delta/(z\phi(z))}$. The descriptive point diagnostic passes iff its supplied reliability $\widehat\lambda_n \ge \lambda^{\dagger}$; under weak information this random plug-in comparison is not itself a population adequacy guarantee. When $\beta_0 = 0$ the population non-centrality vanishes and every reliability is adequate (the population analogue of $t^*_n$ is zero, so $\lambda^{\dagger}=0$). The fixed point is the coherent point evaluation: at the boundary $\widehat\lambda_n = \lambda^{\dagger}$ the corrected pilot $\widehat\beta^*_K/\lambda^{\dagger}$ is evaluated at exactly the reliability used in the comparison. Two practical properties follow from the closed form: $\lambda^{\dagger}$ is computed from the \emph{printed regression output alone} --- indeed from the reported $t$-statistic alone --- with no noise pilot, and it is increasing in $|t^*_n|$: the more precisely a specification claims to estimate a large standardized effect, the more reliability that claim requires. This is the headline number of the empirical section (Table~\ref{tab-vdem}), where we report it against known reliabilities. \emph{Terminology (used without exception in what follows).} The paper reports two distinct objects and never uses one word for both.
\begin{itemize}[leftmargin=1.4em,itemsep=1pt,topsep=2pt]
\item A specification obtains a \textbf{point pass} when $\widehat\lambda_n \ge \lambda^{\dagger}$ with the pilot evaluated at its point estimate. This is a \emph{descriptive} verdict: it says the plug-in diagnostic clears the threshold, and it carries no coverage guarantee, because it ignores the sampling uncertainty of the coefficient pilot.
\item A specification is \textbf{certified} when it obtains a point pass \emph{and} still passes with the coefficient pilot replaced by its upper confidence bound (Remark~\ref{rem-corr-noise}, protocol step~6). This is the conservative, size-controlled object, and it is the only one the word ``certified'' is used for.
\end{itemize}
Every verdict in every table, figure and paragraph below is labelled as one or the other; a point pass is never reported as a certificate, and where a specification is a point pass but not a certificate we say so explicitly. The paper's title refers to the second object, which is the procedure's target; most of the empirical illustrations report the first, which is what the available pilots support.
\end{definition}
The primitive threshold $\tau^2_{\mathrm{crit}}$ cannot have a universal
Stock--Yogo lookup: it depends on $(\rho,c^2,|\beta_0|/\sigma)$. The feasible
fixed point is different. Once the reported $|t^*_n|$ is given, $\rho$ and the
other primitives have already been absorbed into that statistic, so
$\lambda^{\dagger}$ has the compact lookup in Table~\ref{tab-breakdown-lookup}.
\begin{table}[ht]
\centering
\caption{Breakdown reliability from the reported $t$-statistic
($\alpha=\delta=0.05$, $\eta^{\dagger}=0.652$).}
\label{tab-breakdown-lookup}
\begin{tabular}{lrrrrrrr}
\toprule
$|t^*_n|$ & $0.5$ & $1.0$ & $1.96$ & $2.0$ & $3.0$ & $5.0$ & $10.0$ \\
\midrule
$\lambda^{\dagger}$ & $0.434$ & $0.605$ & $0.750$ & $0.754$ & $0.821$ & $0.885$ & $0.939$ \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{0.93\textwidth}\footnotesize
Notes: $\lambda^{\dagger}=|t^*_n|/(|t^*_n|+0.652)$. There is no separate
$\rho$ row because, conditional on the printed $t$-statistic, the fixed-point
threshold is exactly $\rho$-free; $\rho$ remains present in the primitive
$\tau^2$ threshold and in the statistic's sampling law.
\end{minipage}
\end{table}
\begin{proposition}[Exact-normal verdict margin]\label{prop-verdict-margin}
Fix $(\alpha,\delta)$ and write $z=z_{1-\alpha/2}$ and
\[
s(r):=\Phi(-z-r)+1-\Phi(z-r),\qquad r\ge0.
\]
Then: (i) $s(r)$ is strictly increasing for $r>0$, with
$s(0)=\alpha$ and $s(\eta^{\dagger})=\alpha+\delta$; (ii) at any supplied
reliability $\ell\in(0,1]$, the corrected-pilot point diagnostic satisfies
\[
\widehat r(\ell):=|t^*_n|\frac{1-\ell}{\ell}le\eta^{\dagger}
\quad\Longleftrightarrow\quad
\ell\ge\lambda^{\dagger};
\]
and (iii) if an alternative calibration produces a non-centrality magnitude
$\widetilde r$, its pass/fail verdict agrees with that based on $\widehat r$
whenever
\[
|\widetilde r-\widehat r|<|\widehat r-\eta^{\dagger}|.
\]
Thus $|\widehat r-\eta^{\dagger}|$ is an explicit robustness margin for the
binary verdict within the non-central limit experiment, even where the
percentage $s(\widehat r)$ is too far into the tail to be usefully calibrated.
\end{proposition}
\begin{proof}
For $r>0$, $s'(r)=\phi(z-r)-\phi(z+r)>0$, because
$|z-r|<z+r$ and the standard-normal density is strictly decreasing in its
absolute argument. Part (i) follows with the definition of
$\eta^{\dagger}$. For (ii), multiply
$|t^*_n|(1-\ell)/\ell\le\eta^{\dagger}$ by $\ell>0$ and rearrange to obtain
$\ell\ge |t^*_n|/(|t^*_n|+\eta^{\dagger})$. Part (iii) follows because the
stated open ball cannot cross the threshold $\eta^{\dagger}$.
\end{proof}
\begin{remark}[Clustered and serially dependent regression errors]\label{rem-cluster}
Applied panel practice clusters standard errors by unit \citep{Arellano1987, BDM}; see \citet{CameronMiller2015} and \citet{MNW2023} for current practice. The cluster-robust version of the diagnostic is not a heuristic rescaling but a theorem: Section~\ref{sec-cluster} establishes a cluster-level score CLT (Lemma~\ref{lem-cluster-clt}), consistency of the Arellano variance estimator in the many-fixed-effect regime (Lemma~\ref{lem-crve}), and the limit $T^{CR}_n(\beta_0)\Rightarrow N(\eta/\sqrt{\psi},1)$ (Theorem~\ref{thm-cluster}). Clustering enters through the denominator alone, because the attenuation bias in the numerator of $\eta$ is a functional of $\operatorname{\mathbb E}[\nu'M_{K_n}\nu]$ and $\nu\perp u$ by Assumption~\ref{ass-nu}(i); the feasible consequence is
\[
|\eta_{CR}| \;=\; \frac{|\eta|}{\sqrt{\psi}}, \qquad
|\widehat\eta_{CR}| \;=\; \frac{|\widehat\eta|}{\sqrt{\widehat\psi}}, \qquad
\lambda^{\dagger}_{CR} \;=\; \frac{t^*_n/\sqrt{\widehat\psi}}{\,t^*_n/\sqrt{\widehat\psi} + \eta^{\dagger}(\alpha,\delta)\,},
\]
with $\widehat\psi = \tau^{*2}_n\widehat V_{CR}/\widehat\sigma^{*2}_{\mathrm{CJN}}$ the reported variance-inflation factor. The first display is the population target and the second the computed diagnostic; they are separate objects, since $\widehat\psi$ estimates $\psi\sigma^2/\varsigma^2$ rather than $\psi$ (Corollary~\ref{cor-cluster-feasible}(a), Remark~\ref{rem-psi-sign}). Corollary~\ref{cor-cluster-feasible} shows this recipe is legitimate even though, under cluster dependence, neither $\widehat\psi$ nor the feasible $|\widehat\eta|$ is individually consistent for its i.i.d.\ target: the two inconsistencies are the same factor and cancel. Equivalently, the cluster-robust diagnostic is the i.i.d.\ diagnostic run on the reported cluster-robust $t$-statistic.
Three conditions must be checked, and all three are cheap. (1) \emph{Many clusters}: $G_n\to\infty$ with no dominant cluster; the few-clusters regime is not covered (Remark~\ref{rem-cluster-scope}). (2) \emph{Projection compatibility}: fixed effects nested within clusters cost nothing, while Lemma~\ref{lem-nest}(b) supplies the coarse sufficient condition $d^{\mathrm{ne}}_n/G_n\to0$ for a balanced non-nested block. The direct sample analogue is more informative:
\[
\widehat\chi_{\mathrm{proj}}
:=\frac{\sum_g\|Ma^{(g)}-a^{(g)}\|^2}{\tau^{*2}_n}.
\]
Although $d^{\mathrm{ne}}_n/G_n$ is $59/163\approx0.36$ for V-Dem against $7/595\approx0.012$ for PSID, direct computation gives $\widehat\chi_{\mathrm{proj}}=0.0069$--$0.0073$ across the three static V-Dem treatments and $0.00168$ for PSID. The realized loadings therefore satisfy the compatibility requirement closely in both applications; this finite-panel diagnostic supports the reported cluster columns but, by itself, does not prove the asymptotic sequence condition. (3) \emph{No degrees-of-freedom correction}: the conventional $\tfrac{n-1}{n-K_n}$ factor inflates the CRVE by an asymptotic factor $1/(1-\rho)$ and must be omitted (Lemma~\ref{lem-crve}).
Finally, the \emph{direction} of the adjustment is an empirical matter and must be read off $\widehat\psi$ rather than assumed: clustering deflates the measured distortion only when the regressor is persistent as well as the error, and otherwise amplifies it (Remark~\ref{rem-psi-sign}). It happens to deflate it in both applications, where $\widehat\psi$ runs from $2.33$ to $25.2$.
\end{remark}
\subsection{Protocol}\label{sec-protocol}
\begin{framed}
\noindent\textbf{Diagnostic protocol.} Given a saturated FE regression of $Y$ on observed $X^*$:
\begin{enumerate}[label=\arabic*., nosep, start=0]
\item \emph{Scope check.} Verify that $X^*$ is a continuous regressor whose measurement error is plausibly classical (Assumption~\ref{ass-nu}). If $X$ is binary, \textbf{stop}: misclassification error is nonclassical by construction and the diagnostic does not apply (Section~\ref{sec-scope}).
\item Record $\widehat\rho_n = d_{K_n}/n$, $\tau^{*2}_n = X^{*\prime} M_{K_n} X^*$, $\widehat\sigma^*_{\mathrm{CJN}}$, and $\widehat\beta^*_K$ from the regression output.
\item Obtain a noise pilot: $\widehat\sigma_\nu^2$ from a measurement model, a validation study, or a reliability grid (Remark~\ref{rem-est}); set $\widehat a_n = \widehat\sigma_\nu^2 (n - d_{K_n})$ and $\widehat\lambda_n = 1 - \widehat a_n / \tau^{*2}_n$. If $\widehat\lambda_n\le0$, flag the specification and stop; the corrected pilot and certificate require positive reliability.
\item Correct the coefficient pilot: $\widehat\beta_0^{\mathrm{corr}} = \widehat\beta^*_K / \widehat\lambda_n$ (Proposition~\ref{prop-pilot}).
\item \emph{Descriptive point verdict.} Compute the feasible non-centrality $|\widehat\eta| = (|\widehat\beta_0^{\mathrm{corr}}|/\widehat\sigma^*_{\mathrm{CJN}})\,(1-\widehat\lambda_n)\sqrt{\tau^{*2}_n}$ and compare it to the exact-inversion root $\eta^{\dagger}(\alpha,\delta)$ (or the quadratic $\eta^{\mathrm{quad}}=\sqrt{\delta/(z_{1-\alpha/2}\phi(z_{1-\alpha/2}))}$); equivalently compare $\widehat\lambda_n$ to the breakdown $\lambda^{\dagger} = t^*_n/(t^*_n+\eta^{\dagger})$ of Definition~\ref{def-breakdown}, or the distortion $z_{1-\alpha/2}\phi(z_{1-\alpha/2})|\widehat\eta|^2$ to $\delta$, or $\tau^{*2}_n - \widehat a_n$ to $\tau^2_{\mathrm{crit}}$ of~\eqref{eq-cv}. This point verdict is descriptive.
\item \emph{Cluster-robust verdict.} If the reported inference clusters standard errors, form the variance-inflation factor $\widehat\psi$ from the CRVE --- \emph{without} the $\tfrac{n-1}{n-K_n}$ small-sample factor, which over-corrects by $1/(1-\rho)$ (Lemma~\ref{lem-crve}) --- and repeat step~4 with $|\widehat\eta|/\sqrt{\widehat\psi}$ and $\lambda^{\dagger}_{CR}$ (Theorem~\ref{thm-cluster}). First check projection compatibility. The direct route is to evaluate $\sum_g\|Ma^{(g)}-a^{(g)}\|^2/\tau^{*2}_n$, which costs one pass over the residualized regressor; the shortcut is to verify many clusters together with $d^{\mathrm{ne}}_n/G_n$ small, where $d^{\mathrm{ne}}_n$ counts fixed effects \emph{not} nested within clusters, but that reduction is licensed only under the balance conditions (N1)--(N2) of Lemma~\ref{lem-nest}(b) and should not be used without them. Report both verdicts; when they differ, the cluster-robust one is the one that matches the reported standard errors.
\item \emph{Formal certificate (conservative; the only step that yields a certificate).} Because $\widehat\beta_0^{\mathrm{corr}}$ is nondegenerate under weak information --- it is \emph{not} consistent (Corollary~\ref{cor-slope}, Remark~\ref{rem-plugin}) --- replace $|\widehat\beta_0^{\mathrm{corr}}|$ by its upper confidence bound $|\widehat\beta_0^{\mathrm{corr}}| + z_{1-\gamma}\,\widehat{\mathrm{se}}(\widehat\beta_0^{\mathrm{corr}})$, $\widehat{\mathrm{se}}(\widehat\beta_0^{\mathrm{corr}})=\widehat s_{CR}/(\widehat\lambda_n\sqrt{\tau^{*2}_n})$ with $\widehat s_{CR}=\widehat\sigma^*_{\mathrm{CJN}}\sqrt{\widehat\psi}$ (so $\widehat\psi=1$ under i.i.d.\ inference), at a level $\gamma$ named separately from $\delta$, before recomputing $|\widehat\eta|$. Proposition~\ref{prop-certificate} controls the false-certification probability at $\gamma$. Only specifications that pass with this upper bound are \emph{certified}; a point pass that does not survive it must be reported as a point pass only, and never described as certified.
\item Report $\widehat\lambda_n$ (the power tax of Proposition~\ref{prop-power}) and the sensitivity band over $\lambda$ alongside the verdict, stating in every table and figure whether the entry is a point pass, a certificate, or a cluster-rescaled reading. If the specification fails, use the Griliches--Hausman estimators or a fixed-effect Anderson--Rubin confidence set (companion work in progress).
\end{enumerate}
\end{framed}
Unlike the IV case, no universal numerical table for $\tau^2$ can exist. Stock--Yogo's IV thresholds are unit-free constants; here the primitive threshold is data-dependent through $|\beta_0|/\sigma$ and the external noise scale. What \emph{can} be tabulated is the fixed-point reliability as a function of the reported $t$-statistic, as Table~\ref{tab-breakdown-lookup} does. The feasible one-line check~\eqref{eq-feasible} together with Definition~\ref{def-breakdown} then delivers the verdict from regression output plus the noise pilot. When a numerical threshold for $\tau^2$ itself is wanted, it should be read from the exact inversion~\eqref{eq-cv-exact}; the quadratic closed form~\eqref{eq-cv} is mildly anti-conservative ($\eta^{\dagger} < \eta^{\mathrm{quad}}$), with exact-to-quadratic threshold ratios from about $1.01$ far from the vanishing boundary to $1.34$ near it.
\section{Simulations}\label{sec-sim}
Four Monte Carlo designs, in increasing order of importance for the paper's credibility; the two core validations (Designs~1 and~4) are reported here and two supporting checks (Designs~2 and~3) in \ref{app-sim}. All are two-way FE DGPs (units $i=1,\dots,N$, periods $t=1,\dots,T$, $d_K=N+T-1$, $\rho_n=(N+T-1)/(NT)$); fixed effects are removed by the within transformation, so we generate $X_{it}=s\,\varepsilon_{it}$ with $s=\sqrt{\tau^2/n}$, making the primitive $nQ_{K_n}=\tau^2$ hit the target grid exactly, $u_{it}\sim N(0,\sigma^2)$, and $\nu_{it}\sim N(0,\sigma_\nu^2)$. Under this calibration Lemma~\ref{lem-clt} predicts $\operatorname{\mathbb E}[X'M_{K_n}X]=(1-\rho_n)\tau^2$, \emph{not} $\tau^2$ --- a prediction the simulations test directly (Design~1). We state the calibration convention explicitly because the alternative, $s^2=\tau^2/(n-d_K)$ (which holds $\operatorname{\mathbb E}[X'M_{K_n}X]=\tau^2$), makes the $(1-\rho)$ discount invisible to the simulation and so cannot adjudicate it. Seeds are fixed and the Julia scripts are released with the paper. Reported table numbers use $4{,}000$--$5{,}000$ replications per cell; the calibration figure (Figure~\ref{fig-design1}) and Design~1 use $6{,}000$.
\subsection{Design 1 (calibration)} With $\sigma_\nu^2=c^2/n$, we verify that the empirical mean of $T_n^{\mathrm{CJN}*}$ tracks the $\eta_n$ of Theorem~\ref{thm-noncentral} and that empirical size matches the exact non-central prediction $\Phi(-z-\eta_n)+1-\Phi(z-\eta_n)$ across the $(\tau^2,c^2)$ grid. At $N=200$, $T=5$ ($\rho_n=0.204$), $|\beta_0|/\sigma=1$: for $(\tau^2,c^2)=(1,5)$, $\eta_n=-1.82$, mean $T = -1.79$, size $0.433$ (predicted $0.445$); for $(5,1)$, $\eta_n=-0.36$, mean $T=-0.36$, size $0.069$ (predicted $0.065$); for $(10,5)$, $\eta_n=-1.15$, mean $T=-1.17$, size $0.218$ (predicted $0.210$). The empirical within reliability matches the $\rho$-free $\lambda_n$ of~\eqref{eq-lambda} across the grid (maximum deviation $0.003$), and the empirical Hessian ratio $X'M_{K_n}X/(nQ_{K_n})$ sits on $1-\rho_n=0.796$ in every cell --- squarely on the limit of Lemma~\ref{lem-clt}, and far from the naive value $1$. The empirical mean and rejection frequency of $T_n^{\mathrm{CJN}*}$ match the $N(\eta_n,1)$ prediction across the grid (Figure~\ref{fig-design1}), confirming Theorem~\ref{thm-noncentral}.
\begin{figure}[htbp]\centering
\includegraphics[width=0.86\linewidth]{fig_design1.pdf}
\caption{Design 1 calibration ($N=200$, $T=5$, $6{,}000$ replications per cell over $\tau^2\in\{1,2,5,10\}$, $c^2\in\{0.5,1,2,5\}$). (a) empirical mean of the CJN $t$-statistic against the theoretical non-centrality $\eta_n$; (b) empirical size against the exact non-central prediction. Points lie on the $45^\circ$ line (dotted), confirming $T_n^{\mathrm{CJN}*}\Rightarrow N(\eta_n,1)$.}\label{fig-design1}
\end{figure}
\subsection{Designs 2 and 3 (appendix)} Two supporting checks are reported in full in Appendix~\ref{app-sim}. Design~2 (threshold accuracy) sweeps $\tau^2$ and locates the empirical size-crossing: the corrected closed-form threshold~\eqref{eq-cv} matches the empirical threshold to within Monte Carlo accuracy (about $5\%$ on average, at most $9\%$ in any cell) with no systematic dependence on $(N,T,c^2)$, whereas the discarded linear surrogate over-demands the residual signal by a factor of $2.5$--$2.8$ and is anti-conservative at larger $\delta$. Design~3 (pilot sensitivity) confirms Proposition~\ref{prop-pilot}(i) on simulated data --- the naive attenuated pilot passes the point diagnostic for a specification whose true size is $16\%$, which the corrected pilot correctly flags --- and shows that misspecifying the noise pilot by a factor $k\in\{0.5,\dots,2\}$ maps monotonically and boundedly into the distortion estimate, motivating the breakdown-reliability band (Definition~\ref{def-breakdown}).
\subsection{Design 4 (the finite-$n$ mapping check)} For each cell we fix a finite, non-drifting numerical $\sigma_\nu^2$ (a single constant per cell, not an asymptotic sequence) chosen to hit a target within reliability $\lambda_n\in\{0.95,0.9,0.8,0.65,0.5,0.3\}$, generate finite-sample data, and compare the exact empirical size to both the quadratic prediction $\alpha+z_{1-\alpha/2}\phi(z_{1-\alpha/2})\eta_n^2$ and the exact non-central prediction, with $\eta_n$ read off the \emph{finite-$n$} mapping~\eqref{eq-eta-n} (no drifting sequence appears in that mapping). The result (Figure~\ref{fig-design4} and Table~\ref{tab-design4}) is decisive: the exact non-central prediction matches empirical size at \emph{every} $\lambda_n$, and the size depends on the design only through $\lambda_n$. Reproducing the \emph{same} $\lambda_n$ at $n=1000$ and $n=5000$ requires $\sigma_\nu^2=\tau^2(1-\lambda_n)/(\lambda_n n)=O(1/n)$, which is exactly the local drift, so the coincidence across $n$ is not a claim of robustness to a genuinely fixed $\sigma_\nu^2$ as $n\to\infty$ (under which $\lambda_n\to0$ and size $\to1$); it confirms that $\lambda_n$ is the sufficient statistic the drift makes it, and that the finite-$n$ non-central mapping is accurate at fixed finite noise. The drift is thus an \emph{approximation device}, not an artifact of the asymptotics. The quadratic prediction is accurate for $\lambda_n\ge0.8$, under-predicts mildly in the band $\lambda_n\in[0.5,0.65]$ (by at most $1.5$ percentage points), and over-predicts in the far tail ($\lambda_n=0.3$); once $\eta_n$ is large the exact inversion~\eqref{eq-cv-exact} should be used, the behavior anticipated in Section~\ref{sec-drift}. This is the paper's referee-facing figure and belongs in the main text.
\begin{figure}[htbp]\centering
\includegraphics[width=0.66\linewidth]{fig_design4.pdf}
\caption{Design 4 (fixed-noise validation). Size of the nominal-$5\%$ CJN $t$-test under fixed finite-$n$ noise (a single $\sigma_\nu^2$ per cell), against within reliability $\lambda_n$ (worsening to the right). The exact non-central prediction (solid) tracks the empirical size at both $n=1000$ and $n=5000$, which coincide, so size depends on the design through $\lambda_n$ alone. The quadratic prediction $\alpha+z\phi(z)\eta_n^2$ (dashed) tracks the exact curve to within $1.5$ percentage points in the moderate regime ($\lambda_n\ge0.65$, shaded) and departs from it once attenuation is severe (under-predicting near $\lambda_n=0.5$, over-predicting in the far tail). The drift is an approximation device, not an artifact.}\label{fig-design4}
\end{figure}
\begin{table}[htbp]\centering
\caption{Design 4. Exact finite-sample size of the nominal-$5\%$ CJN $t$-test under fixed finite-$n$ noise (a single $\sigma_\nu^2$ per cell) at target within-reliability $\lambda_n$, versus the quadratic and exact non-central predictions. DGP: two-way FE, $\tau^2=5$, $|\beta_0|/\sigma=1$; $5{,}000$ replications. The exact prediction tracks empirical size to within about one percentage point at every $\lambda_n$; size is invariant to $n$ at fixed $\lambda_n$. The quadratic prediction matches through $\lambda_n=0.8$, under-predicts mildly at $\lambda_n\in\{0.65,0.5\}$, and over-predicts in the far tail.}\label{tab-design4}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{lcccc}
\toprule
$\lambda_n$ & $n$ & empirical size & quadratic pred. & exact pred. \\
\midrule
$0.95$ & $1000$ & $0.053$ & $0.051$ & $0.051$ \\
$0.80$ & $1000$ & $0.071$ & $0.073$ & $0.073$ \\
$0.65$ & $5000$ & $0.149$ & $0.136$ & $0.140$ \\
$0.50$ & $5000$ & $0.290$ & $0.279$ & $0.293$ \\
$0.30$ & $5000$ & $0.726$ & $0.798$ & $0.724$ \\
\bottomrule
\end{tabular}
\end{adjustbox}
\end{table}
\subsection{Design 5 (cluster-robust validation)}\label{sec-design5}
The cluster theory of Section~\ref{sec-cluster} is validated on a unit-and-time FE design with $G$ units (which are also the clusters), $T$ periods, within-unit AR(1) errors, and clustering by unit, so that unit effects are nested in clusters and the $T-1$ time effects are not. Six predictions are checked; the script is released with the paper (\texttt{verify\_cluster\_clt.py}), and all reported figures are from $2{,}500$--$4{,}000$ replications.
\emph{(i) The limit itself.} At $G=200$, $T=6$, $\rho_n=0.167$, the empirical mean of $T^{CR}_n$ is $-1.25$ against the predicted $\eta/\sqrt{\psi}=-1.21$, with empirical standard deviation $1.03$; adding the time effects ($d^{\mathrm{ne}}_n/G_n=0.025$) gives $-1.22$ against $-1.21$ with standard deviation $1.02$. Tested in its exact conditional form --- $T^{CR}_n$ centred at $-\beta_0(X^{*\prime}M\nu)/\sqrt{\Psi_n}$ --- the match is within Monte Carlo error at every configuration we ran, including strongly persistent designs.
\emph{(ii) CRVE consistency, and its rate.} The ratio $\widehat V^{\mathrm{sc}}_{CR}/\Psi_n$ of Lemma~\ref{lem-crve} is $0.981$ at $G=50$, $0.994$ at $G=200$ and $0.998$ at $G=800$: the bias is $O(1/G_n)$ and vanishes, as claimed.
\emph{(iii) Projection compatibility is necessary, not decorative.} Pushing the non-nested block to $d^{\mathrm{ne}}_n/G_n=3.25$ ($G=12$, $T=40$) breaks the conclusion exactly as Lemma~\ref{lem-nest}(b) predicts: $\widehat V^{\mathrm{sc}}_{CR}/\Psi_n$ falls to $0.82$ and the standard deviation of $T^{CR}_n$ rises to $1.28$, so the cluster-robust test over-rejects. This is the condition that distinguishes the two applications of Section~\ref{sec-app}.
\emph{(iv) The degrees-of-freedom trap.} Applying the conventional $\tfrac{n-1}{n-K_n}$ small-sample factor at $\rho_n=0.167$ inflates $\widehat V^{\mathrm{sc}}_{CR}/\Psi_n$ to $1.203$, against the predicted over-correction factor $1/(1-\rho_n)=1.200$ --- agreement to three digits. The factor must be omitted; it is not a harmless conservatism, and at the saturation levels common in worker--firm designs it would be substantial.
\emph{(v) The $\sigma$-cancellation of Corollary~\ref{cor-cluster-feasible}.} Under cluster dependence $\widehat\sigma^*_{\mathrm{CJN}}$ converges to $\varsigma=0.80\ne\sigma=1$, and $\widehat\psi$ converges to $\psi\sigma^2/\varsigma^2$ rather than to $\psi$; both ingredients of the feasible recipe are individually inconsistent. Their ratio is not: the recipe returns $-1.22$ against the target $-1.21$. This is the step that licenses computing the diagnostic from i.i.d.\ regression output and dividing by the reported inflation factor.
\emph{(vi) The sign of $\psi-1$.} Three configurations at $G=200$ --- $(\varrho_x,\varrho_u,T) = (0,\,0.6,\,10)$, $(0.9,\,0.9,\,20)$ and $(0.98,\,0.95,\,40)$ --- give $\psi = 0.75$, $1.27$ and $1.72$: clustering amplifies the measured distortion in the serially independent design and deflates it in the persistent ones, as Remark~\ref{rem-psi-sign} describes. These are three points in the design space, not a ceteris-paribus comparative static in $\varrho_x$: the error persistence and the panel length move with the treatment persistence across the three, and $\psi$ is itself increasing in $T$ at fixed $(\varrho_x,\varrho_u)$. The crossing of one is therefore a joint property of persistence and panel length, not a function of $\varrho_x$ alone.
\section{Applications}\label{sec-app}
\subsection{Democracy and growth}\label{sec-app-vdem}
The ideal application has three properties: the regression is canonical and FE-saturated; the regressor is known to be noisily measured; and, rarest of all, the noise variance is externally observable rather than assumed. Country-year democracy--growth panels have all three. The specification of \citet{ANRR} regresses log GDP per capita on a democracy measure with country and year fixed effects (plus lags); the democracy regressor is constructed from expert-coded indices whose disagreement is well documented, and \citet{ANRR} themselves treat measurement error as a first-order concern, addressing it by instrumenting. Crucially, the V-Dem polyarchy index and the Unified Democracy Scores \citep{PMM} publish observation-level posterior standard deviations from an explicit measurement model: $\widehat\sigma_{\nu, it}$ is \emph{data}, not a calibration.
\emph{The within reliability is invariant to the outcome equation; the verdict is not.} The within reliability $\widehat\lambda_n=1-\widehat a_n/\tau^{*2}_n$ depends only on the regressor's measurement model and the fixed-effect design, so the same $\widehat\lambda_n$ applies to the static regression we run, to the dynamic specification of \citet{ANRR}, and to any other regression built on the same measure and FE structure. The verdict does not: it compares $\widehat\lambda_n$ to a breakdown $\lambda^{\dagger}$ that depends on the outcome-specific $t$-statistic and must be recomputed per specification (and Theorem~\ref{thm-noncentral}, derived for the static within regression, does not automatically transfer to dynamic GMM). Since the harder-won ingredient is the one that transfers, a reader can carry $\widehat\lambda_n$ to their own growth equation and recompute the one-line breakdown there; we report a transparent static regression to make that reading as clean as possible.
The exercise: run the country-and-year FE regression of log GDP per capita (Maddison Project 2020) on a continuous V-Dem index as the treatment, 1960--2018; record $\widehat\rho_n$, $\tau^{*2}_n$, $\widehat\beta^*$, $\widehat\sigma^*_{\mathrm{CJN}}$; set $\widehat a_n = \overline{\widehat\sigma^2_{\nu, it}}\,(n - d_{K_n})$ from the published posterior standard deviations, recovered as $\widehat\sigma_{\nu,it}=(\text{codehigh}_{it}-\text{codelow}_{it})/2$; and run the protocol of Section~\ref{sec-protocol}.\footnote{The adequacy diagnostics in this section are computed with the open-source \textsf{PanelAdequacy} software \citep{PanelAdequacySoftware} (the \texttt{eiv\_adequacy} routine); every within-reliability, breakdown, and i.i.d.\ implied-size entry in Tables~\ref{tab-vdem} and~\ref{tab-psid} reproduces from the V-Dem and PSID panels bundled with the package (see the Data availability statement). The cluster-robust columns ($\widehat\psi$, size$_{CR}$) are produced by the supplementary replication archive. One convention matters and is worth stating: Lemma~\ref{lem-crve} shows the CRVE must be formed \emph{without} the $\tfrac{n-1}{n-K_n}$ small-sample factor, which inflates it by an asymptotic $1/(1-\rho)$. Applying that factor would raise the reported $\widehat\psi$ by $2.5\%$ in the V-Dem panel ($\widehat\rho_n=0.025$) and by $16.9\%$ in the PSID panel ($\widehat\rho_n=0.144$); in the latter case $\widehat\psi$ would move from $2.33$ to about $2.72$ and $\lambda^{\dagger}_{CR}$ from $0.61$ to about $0.59$. No verdict in either table turns on the difference, but the PSID level row passes by a correspondingly wider margin, and users reproducing these numbers with default software settings (which apply the factor) should expect the discrepancy.}
The result, in Table~\ref{tab-vdem} and Figure~\ref{fig-vdem}, is a two-pole contrast within a single dataset and a single estimator, with an instructive middle case that only the cluster-robust rescaling exposes. The aggregate polyarchy index obtains a \emph{point pass}: within reliability $\widehat\lambda_n=0.90$, above its breakdown reliability of $0.76$, implied size $5.6\%$ under i.i.d.\ standard errors and $5.0\%$ under the country-clustered standard errors an applied author would report. At the opposite pole, judicial constraints (a component on which expert coders disagree most) is \emph{flagged}: $\widehat\lambda_n=0.41$ against a breakdown of $0.93$, a shortfall of $0.52$ in reliability, with a design-computed non-centrality of $|\eta|\approx12$, nineteen times the breakdown value $\eta^{\dagger}=0.65$. The flag survives the cluster-robust rescaling of Remark~\ref{rem-cluster}, which inflates the standard error by $\sqrt{\widehat\psi}\approx5$ ($\widehat\psi=25.2$) and still leaves $|\eta_{CR}|=2.4$, well above $\eta^{\dagger}$. Legislative constraints is the middle case: flagged under i.i.d.\ standard errors ($\widehat\lambda_n=0.55$ against a breakdown of $0.62$, implied size $14\%$) but obtaining a point pass under the country-clustered standard errors an applied author would actually report ($\widehat\psi=24.0$, $\lambda^{\dagger}_{CR}=0.25$). The direct projection-compatibility ratios are only $0.0069$--$0.0073$ across the three treatments, so the realized design supports applying Theorem~\ref{thm-cluster} despite the conservative rank ratio $d_n^{\mathrm{ne}}/G_n=0.36$. Legislative constraints clears the cluster point threshold by $0.30$ in reliability but does not clear the formal coefficient-uncertainty certificate. Its i.i.d.\ flag is also fragile: Proposition~\ref{prop-equicorr} shows that an equicorrelation share of about $\omega^{*}=0.26$ would overturn it. We therefore report the standardization-dependent middle row without leaning on it substantively. We lean heavily on the judicial-constraints flag, which remains far beyond the threshold after clustering.
\emph{On the implied-size column at large $|\eta|$.} The entries reported as implied sizes are the exact non-central sizes~\eqref{eq-cv-exact} evaluated at the design-computed $\eta$, and they should be read as \emph{threshold comparisons, not as finite-sample size predictions}. The local-drift mapping $\eta \mapsto$ size is calibrated in Design~4 only up to $\lambda_n = 0.30$, where the implied size reaches $0.73$; the judicial-constraints row, at $|\eta| \approx 12$, lies far outside any simulated region, and so does its cluster-robust counterpart at $|\eta_{CR}| = 2.4$. We therefore do not claim that the true finite-sample size of that specification is $100\%$, or that its cluster-robust size is $67\%$. Proposition~\ref{prop-verdict-margin} states the narrower robustness claim that is warranted: with $\eta^{\dagger}=0.652$, the i.i.d.\ magnitude $12.10$ would have to fall by more than $11.45$ (a $94.6\%$ reduction), and the cluster magnitude $2.41$ by more than $1.76$ (a $72.9\%$ reduction), before either binary verdict crosses the exact-normal threshold. Those are robustness margins within the limit experiment, not finite-sample approximation-error bounds. Readers should treat the size column as ordinal beyond roughly $|\eta| \approx 1$, and the same caution applies to the corresponding PSID entries.
The mechanism is transparent: aggregate indices average over many indicators and so average out coding disagreement (small $\widehat\sigma_\nu^2$), whereas the disaggregated constraint sub-indices retain it ($\widehat\sigma_\nu^2$ three to seven times larger). The anti-conservative pilot of Proposition~\ref{prop-pilot} is visible in the table: the naive $|\widehat\beta^*|$ understates the corrected $|\widehat\beta_0^{\mathrm{corr}}|$ by the factor $\widehat\lambda_n$ in every row. The economics is secondary here; the point is that the diagnostic is computable from published uncertainty estimates on a specification the literature cares about, and that it discriminates: one canonical regressor obtains a point pass with margin, a second is flagged under both standardizations, and a third changes from a flag to an uncertified point pass when inference is clustered as applied practice requires.
\begin{table}[htbp]\centering
\caption{Lead application. EIV adequacy diagnostic on a country--year democracy--growth panel (V-Dem regressors, Maddison log GDP per capita, country and year FE, 1960--2018, $N=163$ countries, $n\approx 8.5$--$8.9$k, $\widehat\rho_n\approx0.025$). Posterior noise $\widehat\sigma_{\nu,it}$ from the V-Dem measurement model. ``Breakdown'' is the fixed-point $\lambda^{\dagger}$ of Definition~\ref{def-breakdown} at $\delta=0.05$; a specification obtains a \emph{point pass} iff $\widehat\lambda_n\ge\lambda^{\dagger}$. The ``verdict'' columns report point passes and flags. The formal certificate of Proposition~\ref{prop-certificate} is a separate, stronger object evaluated at the coefficient pilot's upper confidence bound (protocol step~6); of the rows below only aggregate polyarchy clears it, at $\gamma=0.05$ (see the discussion following the table). ``Size'' is the local-drift implied size of the nominal-$5\%$ $t$-test and should be read ordinally beyond $|\eta|\approx1$; it is not a finite-sample size prediction (see the discussion following this table and Design~4, whose calibration stops at implied size $0.73$). The last three columns apply the cluster-robust diagnostic of Theorem~\ref{thm-cluster}: $\widehat\psi$ is the CRVE variance-inflation factor and the verdict compares $\widehat\lambda_n$ to $\lambda^{\dagger}_{CR}$. Although the coarse sufficient-condition ratio is $d^{\mathrm{ne}}_n/G_n=59/163\approx0.36$, the direct realized projection ratios $\widehat\chi_{\mathrm{proj}}$ are $0.0069$--$0.0073$; this supports the cluster columns for this design but is not a proof of the asymptotic sequence condition.}\label{tab-vdem}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{lcccccccc}
\toprule
& \multicolumn{2}{c}{pilots} & \multicolumn{3}{c}{i.i.d.\ standard errors (theory)} & \multicolumn{3}{c}{clustered by country (rescaling)} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-6}\cmidrule(lr){7-9}
V-Dem treatment & $\widehat\lambda_n$ & $|\widehat\beta^*|\,/\,|\widehat\beta_0^{\mathrm{corr}}|$ & $\lambda^{\dagger}$ & size & verdict & $\widehat\psi$ & size$_{CR}$ & verdict$_{CR}$ \\
\midrule
Polyarchy (aggregate) & $0.90$ & $0.061\,/\,0.068$ & $0.76$ & $5.6\%$ & point pass & $19.2$ & $5.0\%$ & point pass \\
Legislative constraints & $0.55$ & $0.024\,/\,0.045$ & $0.62$ & $14\%$ & \textbf{flagged} & $24.0$ & $5.4\%$ & point pass\textsuperscript{$\S$} \\
Judicial constraints & $0.41$ & $0.224\,/\,0.543$ & $0.93$ & $\ge0.99$ & \textbf{flagged} & $25.2$ & $\approx67\%$ & \textbf{flagged} \\
\bottomrule
\end{tabular}
\end{adjustbox}
\par\smallskip
\begin{minipage}{0.95\textwidth}\footnotesize\raggedright
\textsuperscript{$\S$} Point pass under the cluster-robust diagnostic, flagged under the i.i.d.\ one. Since inference in this literature clusters by country, the cluster-robust column is operative; the point pass is not a formal certificate. Proposition~\ref{prop-equicorr} independently shows that the i.i.d.\ flag is sensitive to persistent coder error.
\end{minipage}
\end{table}
\begin{figure}[htbp]\centering
\includegraphics[width=0.72\linewidth]{fig_vdem.pdf}
\caption{The two-pole diagnostic on the V-Dem panel. For each treatment, the filled point is the actual within reliability $\widehat\lambda_n$ and the vertical tick is the breakdown reliability $\lambda^{\dagger}$ (Definition~\ref{def-breakdown}, $\delta=0.05$). A specification obtains a \emph{point pass} iff $\widehat\lambda_n\ge\lambda^{\dagger}$ --- a descriptive verdict, not a formal certificate: the aggregate polyarchy index clears its breakdown, the institutional-constraint sub-indices fall below theirs.}\label{fig-vdem}
\end{figure}
Three caveats accompany the table. \emph{First}, expert-coding error may be serially dependent within country, violating the i.i.d.\ assumption on $\nu$. The equicorrelated extension of Section~\ref{sec-ext} covers this, and its direction is one-sided: the marginal-variance diagnostic \emph{overstates} the noise, so a point pass survives a fortiori under any within-country persistence, while a flag is conclusive only given a bound on the common-component share $\omega$. Proposition~\ref{prop-equicorr} makes the bound exact. Writing $\widehat\lambda_n$ for the reliability under the i.i.d.\ benchmark and $\lambda_\omega$ for the effective reliability, $(1-\lambda_\omega)/\lambda_\omega=(1-\omega)(1-\widehat\lambda_n)/\widehat\lambda_n$, so a baseline flag is overturned when
\[
\omega\ge\omega^*:=1-\frac{\widehat\lambda_n}{1-\widehat\lambda_n}\frac{1-\lambda^\dagger}{\lambda^\dagger}.
\]
This gives $\omega^*\approx0.95$ for judicial constraints, so that flag stands unless coder error is almost entirely common within country, against $\omega^*\approx0.25$ for legislative constraints, which is one of the two reasons we lean on that row only lightly. \emph{Second}, the specification is static two-way FE; adding lagged output in the manner of \citet{ANRR} leaves the two-pole contrast intact but changes the magnitudes, since a lagged dependent variable introduces an incidental-parameter bias of different origin that we do not conflate with the measurement-error channel. \emph{Third}, the verdicts in Table~\ref{tab-vdem} are point passes and flags, but the aggregate polyarchy row also clears the \emph{formal certificate} of Proposition~\ref{prop-certificate}, which is the size-controlled object. Apart from the already supplied reliability $\widehat\lambda_n$, the certificate uses regression output: $\widehat{\mathrm{se}}(\widehat\beta_0^{\mathrm{corr}})=\widehat s_{CR}/(\widehat\lambda_n\sqrt{\tau^{*2}_n})$. At $\gamma=0.05$ the upper bound is $U=0.122$ against $|\widehat\beta_0^{\mathrm{corr}}|=0.068$, giving $\widehat\eta^{\,U}=0.41$ against $\eta^{\dagger}=0.65$: polyarchy is certified, with the false-certification probability controlled at $5\%$, and it remains certified under the country-clustered scale ($\widehat\eta^{\,U}=0.24$). Judicial constraints fails the point diagnostic by $0.52$ in reliability and a fortiori fails the certificate. These figures are recovered from the quantities printed in Table~\ref{tab-vdem} rather than recomputed from the raw panel; the polyarchy verdict is unchanged when the implied $t^*_n$ is perturbed by $\pm20\%$ (\texttt{verification/certificate\_calc.py}), so it does not rest on the rounding.
\subsection{Mechanism illustration: survey earnings}\label{sec-app-psid}
A complementary mechanism illustration uses the canonical noisy regressor of the validation literature, self-reported earnings, and illustrates source (ii) of Remark~\ref{rem-est}: the noise pilot is an \emph{external} reliability ratio, not a per-observation posterior. It isolates a second mechanism by which the threshold binds, distinct from the first application. There, two poles arose from different regressors (an aggregate index versus its noisier sub-components). Here a single regressor crosses the threshold under the fixed-effect \emph{transformation} itself: the Griliches--Hausman amplification made quantitative. The exercise is not offered as a preferred substantive labor-supply specification.
We use the Cornwell--Rupert panel (595 individuals over 1976--1982, derived from the PSID and public through the \texttt{plm} package), and run a person-and-year FE regression of weeks worked on the log wage. This deliberately simple equation is useful because the log wage has external validation evidence and the same observations support a transparent level-versus-within comparison; its coefficient is not interpreted as a structural labor-supply parameter. The within transformation is severe: the within standard deviation of the log wage is $0.15$, against a cross-sectional $0.46$, so two-thirds of the signal is removed by the fixed effects, exactly the regime in which \citet{GH} warn that reliability collapses. The regression returns $\widehat\rho_n = 0.14$, $\tau^{*2}_n = 83.2$, an attenuated $\widehat\beta^* = 0.73$ (weeks per log-wage unit) and $\widehat\sigma^*_{\mathrm{CJN}} = 4.25$, giving a breakdown reliability of $\lambda^{\dagger} = 0.71$ at $\delta = 0.05$. This breakdown is \emph{self-consistent}: because the corrected coefficient pilot $\widehat\beta_0^{\mathrm{corr}} = \widehat\beta^*/\lambda$ itself depends on the reliability at which the diagnostic is evaluated, the single reported breakdown solves the fixed point $\lambda = \lambda^{\dagger}(\lambda)$, i.e.\ $\lambda^{\dagger} = 1/(1 + \eta^{\dagger}\widehat\sigma^*_{\mathrm{CJN}}/(|\widehat\beta^*|\sqrt{\tau^{*2}_n})) = 0.71$; evaluating $\lambda^{\dagger}$ instead at any single external reliability leaves the pass/fail verdicts of Table~\ref{tab-psid} unchanged.
The noise pilot comes from the validation literature. \citet{BK}, matching CPS records to Social Security earnings, estimate a cross-sectional reliability of log annual earnings of $0.82$ for men, which falls to $0.65$ in first differences because the measurement error is mean-reverting; \citet{BoundEtAl94}, using the PSID Validation Study, place the within/difference-transformed reliability as low as $0.5$. Applying these external reliabilities to the diagnostic (Table~\ref{tab-psid} and Figure~\ref{fig-psid}) delivers a clean verdict: at the \emph{level} reliability $0.82$ the specification passes the point diagnostic (implied size $6.4\%$), because $0.82$ exceeds the point breakdown $0.71$. That level pass does \emph{not} survive the formal certificate of Proposition~\ref{prop-certificate}: at $\gamma=0.05$ the upper bound gives $\widehat\eta^{\,U}=0.71$ against $\eta^{\dagger}=0.65$, equivalently a certificate breakdown of $\approx0.83$ against $\widehat\lambda_n=0.82$. It fails narrowly --- the verdict flips if the implied $t^*_n$ is $20\%$ smaller --- and we therefore report it as a point pass that is not a certificate, which only reinforces that it is the within transformation, not the level regression, that moves this regressor decisively. Every \emph{within}-transformed reliability in the validated range $[0.50, 0.65]$ falls below the breakdown and flags it (implied size $13.6\%$ to $35\%$). The cluster-robust diagnostic of Theorem~\ref{thm-cluster} moderates but does not eliminate the alarm: the CRVE inflation is $\widehat\psi=2.33$, the rescaled breakdown falls to $\lambda^{\dagger}_{CR}=0.61$, and the top of the validated band ($0.65$) then obtains a point pass (rescaled implied size $8.6\%$) while the rest of the band remains flagged. The fixed-effect transformation moves the same regressor from a point pass to a flag by lowering its reliability across $\lambda^{\dagger}$. This is the Griliches--Hausman amplification expressed as a size distortion, and it is the reason the breakdown reliability, not a level reliability read off a validation table, is the quantity to compare against.
\begin{table}[htbp]\centering
\caption{Mechanism illustration. EIV diagnostic for a person-and-year FE regression of weeks worked on the log wage (Cornwell--Rupert PSID panel, $N=595$, $T=7$, $n=4165$, $\widehat\rho_n=0.14$), using \emph{external} reliability pilots. Self-consistent breakdown reliability $\lambda^{\dagger}=0.71$ at $\delta=0.05$ (Definition~\ref{def-breakdown}); under person-clustered inference (Theorem~\ref{thm-cluster}, $\widehat\psi=2.33$) the breakdown is $\lambda^{\dagger}_{CR}=0.61$. The direct projection ratio is $\widehat\chi_{\mathrm{proj}}=0.00168$ (the coarse ratio is $d^{\mathrm{ne}}_n/G_n=7/595\approx0.012$). The corrected pilot is $\widehat\beta_0^{\mathrm{corr}}=\widehat\beta^*/\lambda$; ``size'' is the local-drift implied size of the nominal-$5\%$ $t$-test and is ordinal beyond $|\eta|\approx1$ (the bottom two rows), not a finite-sample size prediction. No row is a formal certificate. \textsuperscript{$\ddagger$}The level row obtains a point pass but fails the formal upper-bound certificate (its coefficient-uncertainty-inflated breakdown is $\approx0.83>0.82$; see text). \textsuperscript{$\S$}Point pass under person-clustered inference, flagged under i.i.d.\ standard errors.}\label{tab-psid}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{lccccccc}
\toprule
& & \multicolumn{3}{c}{i.i.d.\ standard errors} & \multicolumn{3}{c}{clustered by person} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
reliability $\lambda$ (source) & regime & $|\eta|$ & size & verdict & $|\eta_{CR}|$ & size$_{CR}$ & verdict$_{CR}$ \\
\midrule
$0.82$ \ (\citealp{BK}, men, cross-section) & level & $0.35$ & $6.4\%$ & point pass\textsuperscript{$\ddagger$} & $0.23$ & $5.6\%$ & point pass \\
$0.65$ \ (\citealp{BK}, men, first difference) & within & $0.85$ & $13.6\%$ & \textbf{flagged} & $0.56$ & $8.6\%$ & point pass\textsuperscript{$\S$} \\
$0.55$ \ (\citealp{BoundEtAl94}, PSID-VS) & within & $1.29$ & $25.2\%$ & \textbf{flagged} & $0.84$ & $13.5\%$ & \textbf{flagged} \\
$0.50$ \ (\citealp{BoundEtAl94}, PSID-VS, low) & within & $1.58$ & $35.0\%$ & \textbf{flagged} & $1.03$ & $17.8\%$ & \textbf{flagged} \\
\bottomrule
\end{tabular}
\end{adjustbox}
\end{table}
\begin{figure}[htbp]\centering
\includegraphics[width=0.66\linewidth]{fig_psid.pdf}
\caption{Griliches--Hausman amplification in the PSID application. Implied size of the nominal-$5\%$ $t$-test as a function of the log-wage regressor's within reliability $\lambda$ (worsening to the right), from the observed $\hat\beta^*,\widehat\sigma^*_{\mathrm{CJN}},\tau^{*2}_n$. The fixed-effect transformation lowers reliability from the cross-sectional $0.82$ (a point pass, left of the breakdown $\lambda^{\dagger}=0.71$, though not a formal certificate --- see Table~\ref{tab-psid}, note $\ddagger$) into the validated within band $[0.50,0.65]$ (flagged), driving implied size from $6\%$ to $14$--$35\%$. Implied sizes beyond the breakdown are ordinal readings of the local-drift mapping, not finite-sample size predictions.}\label{fig-psid}
\end{figure}
Two caveats are worth stating. The published reliabilities are for log annual \emph{earnings}, whereas the regressor here is the log hourly wage; the two share the mean-reverting-error property that drives the reliability drop under differencing, but their exact reliabilities differ, so the table is best read as illustrating the mechanism rather than certifying this particular labor-supply coefficient. And the within (all-leads-and-lags) reliability of a fixed-effect estimator is not identically the first-difference reliability; both fall below the level reliability, and the band $[0.50, 0.65]$ brackets the range the validation studies support.
\section{Extensions}\label{sec-ext}
\subsection{Vector treatment} With $X \in \mathbb R^p$, $Q_K = \operatorname{\mathbb E}[\operatorname{Var}(X \mid \mathcal G_K)]$ is a $p \times p$ matrix and $\Sigma_\nu$ the noise covariance under the drift $\Sigma_{\nu, n} = C/n$. The construction of Theorem~\ref{thm-noncentral} carries over: the Wald statistic for $H_0:\beta=\beta_0$ is asymptotically non-central $\chi^2_p$ with non-centrality $\eta'\eta$, where the vector non-centrality is governed by the matrix analogue of $\tau^{*2}=(1-\rho)(\tau^2+c^2)$, and a sufficient condition for leading-order size control at tolerance $\delta$ is a lower bound on $\lambda_{\min}(nQ_{K_n})$, mirroring \citet[\S 6]{SY}. We record this as the natural generalization and do not develop it here: the scalar diagnostic is complete on its own, and a sharp multivariate eigenvalue threshold is a separate exercise whose details we omit.
\subsection{Heteroskedasticity} Under conditional heteroskedasticity the CJN homoskedastic variance estimator is replaced by a leave-one-out (HC2) estimator in the manner of \citet{KSS} and \citet{Jochmans2022}. The bias term $-\beta_0\,c^2(1-\rho)$ is unchanged, since it arises entirely from $\operatorname{\mathbb E}[\nu'M_{K_n}\nu]$, which depends on the measurement-error variance alone and not on the conditional variance structure of $u$; heteroskedasticity changes only the standardization of the score. That standardization, however, is \emph{not} in general the true-regressor limit $\omega$ of Assumption~\ref{ass-reg}(iv): the contaminated score $X^{*\prime}M_{K_n}u$ weights the heterogeneous $\sigma_i^2$ through the \emph{contaminated} leverage $\widetilde X^{*2}_{K_n,i}$ rather than $\widetilde X_{K_n,i}^2$, so its variance has its own limit $\omega_*^2 = \lim (X^{*\prime}M_{K_n}X^*)^{-1}\sum_i \widetilde X^{*2}_{K_n,i}\sigma_i^2$. This normalized limit can itself depend on the saturated design; no general $\rho$-free conclusion follows under heteroskedasticity. The non-centrality retains its form with $\omega_*$ in place of $\sigma$, and $\omega_*$ is what the HC2 estimator consistently targets; it coincides with $\sigma$ under homoskedasticity. For cluster dependence --- serial correlation of $u$ within unit, the empirically dominant departure in panels --- the corresponding rescaling of the diagnostic is Remark~\ref{rem-cluster}, and the applications report its verdicts alongside the i.i.d.\ ones.
\subsection{Serially dependent measurement error} The classical assumption that $\nu$ is i.i.d.\ is violated when expert-coding error persists within country (Section~\ref{sec-app-vdem}) or survey error persists within person. The noise covariance enters the \emph{mean} bias only through $\operatorname{\mathbb E}[\nu'M_{K_n}\nu]=\operatorname{tr}(M_{K_n}\Sigma_\nu)$, so the bias term is $-\beta_0\lim\operatorname{tr}(M_{K_n}\Sigma_{\nu,n})$ for a general $\Sigma_\nu$. The trace alone does not, however, deliver the quadratic-form concentration or the score CLT that the non-centrality theorem also needs; those require additional structure on $\Sigma_\nu$ beyond its trace. The within-unit-equicorrelated case (the empirically relevant one for persistent coding error) admits such structure, a one-factor model, and has a clean, favorable answer.
\begin{proposition}[Equicorrelated measurement error]\label{prop-equicorr}
Under the remaining conditions of Theorem~\ref{thm-noncentral}, suppose the unit dummies are among the fixed effects and the measurement error follows the one-factor model
\[
\nu_{it}=\sigma_{\nu,n}\big(\sqrt{\omega}\,g_i+\sqrt{1-\omega}\,\epsilon_{it}\big),\qquad \omega\in[0,1),\ \ \sigma_{\nu,n}^2=c^2/n,
\]
with $\{g_i\}$ and $\{\epsilon_{it}\}$ mutually independent, mean zero, unit variance, each satisfying Assumption~\ref{ass-nu}(ii) (so $\{\epsilon_{it}\}$ is i.i.d.\ with $r=2$), and both independent of $(X,u)$ given $\mathcal G_{K_n}$. This induces exactly the within-unit-equicorrelated covariance $\Sigma_{\nu,n}=\sigma_{\nu,n}^2[(1-\omega)I_n+\omega B]$, $B_{(it),(js)}=\mathbf 1\{i=j\}$. Then Lemma~\ref{lem-atten}, Lemma~\ref{lem-lev-star}, and Theorem~\ref{thm-noncentral} hold verbatim with $c^2(1-\rho)$ replaced by $(1-\omega)\,c^2(1-\rho)$, giving
\[
\eta_\omega=-\frac{\beta_0\,(1-\omega)\,c^2\sqrt{1-\rho}}{\sigma\sqrt{\tau^2+(1-\omega)\,c^2}},\qquad |\eta_\omega|\le|\eta|\ \text{and non-increasing in }\omega.
\]
\end{proposition}
\begin{proof}
Deferred to \ref{app-equicorr}. The key identity is $M_{K_n}\nu = \sigma_{\nu,n}\sqrt{1-\omega}\,M_{K_n}\epsilon$: because the unit dummies lie in the fixed-effect column space, $M_{K_n}$ annihilates the entire common factor $g$, not merely its mean. Every quadratic and bilinear form in $\nu$ then reduces to the i.i.d.\ idiosyncratic $\epsilon$ with noise scale $(1-\omega)c^2/n$, so Lemma~\ref{lem-atten} and Lemma~\ref{lem-lev-star} apply verbatim with $c^2$ replaced by $(1-\omega)c^2$, giving $\eta_\omega=-\beta_0(1-\omega)c^2\sqrt{1-\rho}/(\sigma\sqrt{\tau^2+(1-\omega)c^2})$; monotonicity in $\omega$ follows since $c^2\mapsto|\beta_0|c^2\sqrt{1-\rho}/(\sigma\sqrt{\tau^2+c^2})$ is increasing in $c^2$ and $(1-\omega)c^2$ is decreasing in $\omega$.
\end{proof}
The unit fixed effect \emph{annihilates} the persistent (common) component of the noise, so persistent coding error produces a \emph{smaller} EIV distortion than the classical i.i.d.\ benchmark. Two consequences for the lead application. First, the diagnostic computed from the marginal posterior variances (which ignore within-country correlation) is \emph{conservative}: it uses $a=c^2(1-\rho)$ in place of the true $(1-\omega)a\le a$, so it overstates the noise and issues too few passes. A specification it passes (polyarchy) passes \emph{a fortiori} under any within-country error persistence. Second, the pass direction is the one that is robust: because $|\eta_\omega|$ is decreasing in $\omega$ with $|\eta_\omega|\to 0$ as $\omega\to 1$, a \emph{flag} computed under the i.i.d.\ benchmark is not immune to serial correlation: sufficiently persistent coding error would eventually rescue any finite i.i.d.\ non-centrality. Writing $\widehat\lambda_n$ for the i.i.d.-benchmark reliability and $\lambda_\omega$ for the effective reliability,
\[
\frac{1-\lambda_\omega}{\lambda_\omega}=(1-\omega)\frac{1-\widehat\lambda_n}{\widehat\lambda_n},
\]
the flag is overturned exactly when $\omega\ge\omega^*:=1-[\widehat\lambda_n/(1-\widehat\lambda_n)][(1-\lambda^\dagger)/\lambda^\dagger]$. This separates the two flags of the lead application: $\omega^*\approx0.95$ for judicial constraints (the flag stands unless coder error is almost entirely a common country-level component) against $\omega^*\approx0.25$ for legislative constraints (the i.i.d.\ flag is more fragile, which is a second and independent reason --- alongside its reversal under the cluster-robust diagnostic of Theorem~\ref{thm-cluster} in Table~\ref{tab-vdem} --- not to lean on that row). I.i.d.-based \emph{failure}, unlike an i.i.d.-based pass, is conclusive only given a bound on $\omega$. This is the same mechanism, and the same $M\mathbf 1_i=0$ identity, that motivates the cluster-robust rescaling of Remark~\ref{rem-cluster}; note that here it is a proved consequence of Proposition~\ref{prop-equicorr}, whereas there it is not.
\subsection{Other bias sources} The derivation pattern (introduce a local bias source, compute the non-centrality of the $t$-statistic under the $(\rho, \tau^2)$ drift, invert the leading size distortion) applies beyond classical EIV: heterogeneous treatment effects in TWFE designs \citep{dCdH}, FE misspecification, omitted nonlinearities, and binary misclassification. Each candidate produces its own threshold structure; developing them is left to future work. The last of them is the subject of the scope warning of Section~\ref{sec-scope}, which we do not repeat here beyond restating its conclusion: the diagnostic of this paper must not be applied to a binary mismeasured treatment.
\section{Discussion}\label{sec-conc}
The weak-identification analogy between fixed-effect saturation and weak instruments is genuine but requires care. Under strict exogeneity and homoskedastic errors, $\tau^2 = nQ_K$ does not behave like an IV concentration parameter --- the FE-OLS estimator is unbiased, and no Stock--Yogo threshold arises in the baseline model. Once a bias source is introduced (classical measurement error in the treatment, in our development), the analogy is restored, with a closed-form critical value~\eqref{eq-cv}, a feasible one-line form~\eqref{eq-feasible}, a corrected plug-in protocol, and a power reading: the within reliability $\lambda$ is simultaneously the size-distortion driver, the local-power tax, and the single number a referee should ask for when a saturated specification leans on a noisy regressor. Under Condition~(B), the FE dimension $\rho$ enters only through an overall $\sqrt{1-\rho}$ asymptotic scaling common to signal and noise, not through the reliability $\lambda$ itself (Remark~\ref{rem-rho}); this is a modest role for saturation, and the feasible protocol never decomposes the diagnostic into $(\rho,\tau^2,c^2)$ in the first place.
Three structural features distinguish the threshold from the IV case. The threshold combines a signal-to-noise ratio $\tau^2/c^2$, exactly as classical attenuation does, with a saturation-driven scale factor $\sqrt{1-\rho}$ that raises the standard error without touching the reliability. The primitive $\tau^2$ threshold is data-dependent through $|\beta_0|$, $\sigma$, and the noise scale, so it cannot be summarized in a universal table the way Stock--Yogo IV thresholds can. The feasible fixed-point reliability \emph{can} be tabulated against the reported $t$-statistic (Table~\ref{tab-breakdown-lookup}), while~\eqref{eq-feasible} reduces the diagnostic to one inequality in reported quantities. Finally, the threshold requires an external noise input; Remark~\ref{rem-est} identifies three practical sources, including expert-coded posteriors, validated survey measures, and administrative matches.
For specifications where the threshold is not met, two remedies exist: the estimation route of \citet{GH}, and a fixed-effect Anderson--Rubin test with uniform validity over $\tau^2 \in (0, \infty]$ --- which we develop in companion work in progress --- giving honest confidence sets even when identification is fragile.
Three limitations should travel with the results. First, the diagnostic assumes \emph{classical} error in a \emph{continuous} regressor; it does not apply to binary-treatment misclassification, where the error is nonclassical by construction, and applying it there is a misuse (Section~\ref{sec-scope}). Second, the paper reports two distinct objects and they should not be conflated: a \emph{point pass} is the descriptive verdict that the plug-in diagnostic clears the breakdown reliability, while a \emph{certificate} additionally survives the coefficient pilot's sampling uncertainty (Definition~\ref{def-breakdown}, protocol step~6). Proposition~\ref{prop-certificate} supplies the certificate and its guarantee: the false-certification probability is controlled at a level $\gamma$ named separately from the size tolerance $\delta$ (Remark~\ref{rem-gamma-delta}). Aggregate V-Dem polyarchy is certified at $\gamma=0.05$; no other empirical row is reported as certified, even where a cluster-rescaled point pass obtains. Third, the cluster-robust theory of Section~\ref{sec-cluster} requires many clusters and an asymptotic projection-compatibility condition. The direct sample ratios are small in both applications ($0.0069$--$0.0073$ for V-Dem and $0.00168$ for PSID), which is reassuring for the realized designs, but a single finite-panel calculation cannot establish the required sequence, and neither application covers the few-clusters regime. Extending the derivation pattern to nonclassical misclassification, and to few-cluster asymptotics, are the two developments that would most enlarge the diagnostic's reach.
\section*{Declaration of competing interest}
The author declares that he has no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
\section*{Funding}
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
\section*{Declaration of generative AI and AI-assisted technologies in the manuscript preparation process}
During the preparation of this work the author used an AI-assisted proof-checker tool to perform supplementary consistency checks on the mathematical arguments (the theorems, lemmas, propositions, and their proofs), in addition to the author's own verification. After using this tool, the author reviewed and edited the content as needed and takes full responsibility for the content of the published article.
\section*{Data availability}
All data used in the applications are publicly available: the V-Dem indices and their measurement-model posterior standard deviations from the V-Dem dataset, GDP per capita from the Maddison Project Database 2020, and the Cornwell--Rupert PSID earnings extract distributed with the \texttt{plm} R package. The diagnostics are implemented in the \textsf{PanelAdequacy} software \citep{PanelAdequacySoftware}, available as a \textsf{Julia} package (\texttt{PanelAdequacy.jl}) and an \textsf{R} package (\texttt{panelcert}); the V-Dem and PSID analysis panels are bundled with both packages, so every number in Section~\ref{sec-app} reproduces from the package alone. A replication archive that reproduces the data construction, simulations, figures, and application numbers without relying on the package accompanies the journal submission.
\bibliographystyle{elsarticle-harv}
\bibliography{bibliography.bib}