EconBase
← Back to paper

Reliable Panel Regression: A Default Workflow for Slow-Moving, Mismeasured Variables

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

48,059 characters · 15 sections · 28 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Reliable Panel Regression A Default Workflow for Slow-Moving, Mismeasured Variables

\setcounter{secnumdepth}{-2}

abstract\begin{singlespace} Political scientists often interpret coefficient shrinkage under fixed effects as evidence that pooled associations are confounded. This paper shows why that inference is unreliable for slow moving, mismeasured regressors. Fixed effects can remove much of the signal and identify coefficients from within unit variation that is disproportionately measurement error, attenuating estimates toward zero. A lone fixed effects coefficient may therefore be unable to distinguish confounding from measurement error. I show that the attenuation depends on a regressor's empirical intraclass correlation and measurement reliability. I then propose a default workflow for panel regression. Researchers estimate reliability when possible, report pooled and fixed effects estimates with corrected within reliability, use partial identification bounds when the estimates share a sign, and report fixed effects as a within unit estimate when they do not. For variables with no reliability estimate, I introduce an autocorrelation frontier that bounds the attenuation factor directly. I conclude by applying this workflow to several published results to show that the data often cannot distinguish attenuation from confounding, and the workflow makes clear which case the researcher faces. \noindentKeywords: measurement error; fixed effects; panel data; partial identification; attenuation bias \end{singlespace}

\thispagestyle{empty} \setcounter{page}{1}

Introduction

Does democracy cause economic growth? Is there a resource curse? Do stronger bureaucracies protect democracy? These are important questions in comparative politics and international relations and scholars typically use one-way or two-way fixed effects models to answer them. Often, a large, significant pooled relationship shrinks once one adds fixed effects. Researchers usually interpret that shrinkage as evidence that the former association was confounded. One then treats the latter estimate as the more credible result because it controls for time-invariant confounders and common shocks.

This paper shows why that interpretation is often wrong. Fixed effects remove between-unit variation and use within unit change to identify coefficients. That is useful when the within unit change is informative, but many variables political scientists care about, including regime type, bureaucratic quality, military capability, trade dependence, national income, and oil rents, move slowly and are measured with error. For these variables, much of the important variation lies between units or in long-run trends. The short run movement left after fixed effects can contain significant measurement error. In that case, fixed effects do not simply remove confounding. They also make the regressor less reliable, so the estimated coefficient is pulled toward zero even when the underlying relationship exists.

As a result, finding shrinkage in fixed effects models is ambiguous. It may mean that the pooled estimate was biased, but it may also mean that fixed effects have removed much of the signal and left the estimate to be identified from low reliability within unit variation. The usual interpretation assumes the first story and can push researchers toward substantive null conclusions: democracy does not cause growth, oil does not weaken accountability, bureaucratic quality does not matter. The more cautious conclusion is that the fixed effects estimate may not identify the effect with the available measurement information.

In this paper, I propose a reporting standard for this problem. The main diagnostic is the corrected within reliability of the regressor, denoted $\lambda_w$. This quantity tells the researcher how much of the variation used by fixed effects is signal rather than measurement error. To calculate it, the researcher needs the empirical intraclass correlation (ICC) of the observed regressor and an estimate or range of estimates for the regressor's measurement reliability. The empirical ICC measures how much of a variable's variation lies between units. Slow moving variables have high ICCs, and fixed effects discard exactly that variation. The reliability term captures how precisely a persistent variable is measured.

I argue that researchers should report pooled OLS, fixed effects, the empirical ICC of the key regressor, and the corrected within reliability. A simple sign test then determines what should be reported. If the estimates share a sign and fixed effects is smaller in magnitude, researchers should report partial identification bounds on the effect. If the estimates have opposite signs, they should report fixed effects as a within unit result. If fixed effects is larger in magnitude, the result is a complement rather than an attenuation problem. For variables with no reliability estimate, I introduce an autocorrelation frontier that asks whether measurement error can plausibly explain the shrinkage using the regressor's own persistence. The Diagnostic section develops each case. I provide an open-source R package, ferobust ferobust2026, that automates the diagnostic.

I apply this reporting standard to the case of the effect of bureaucratic quality on economic growth cornell2020. In this example, the fixed effects estimate is positive but insignificant. Once measurement error is incorporated, however, the identified set excludes zero under a broad range of reliability. The fixed effects null therefore cannot be interpreted as evidence of no within unit effect.

A broader audit of the literature shows that, in most cases, the reporting standard does not overturn a null finding. It either tells researchers to report fixed effects because the pooled and within unit estimates have opposite signs, shows that measurement error is unlikely to explain a null finding, or concludes that the data cannot distinguish attenuation from confounding. The point of this article is not to argue that fixed effects models are usually wrong. Rather, the point is that a fixed effects coefficient by itself does not tell scholars which of these situations they are in. Accordingly, this paper is similar to other recent work that implores quantitative political scientists to be clear-eyed and humble when drawing conclusions from their work arel2026quantitative.

The argument also complements the literature on pathologies in two-way fixed effects designs with staggered binary treatments and heterogeneous effects goodmanbacon2021,dechaisemartin2020,callaway2021,sunandabraham2021. That literature addresses difference-in-differences designs with binary treatments. But many panel studies in political science instead use continuous, slow-moving regressors measured across countries or organizations over time, which it does not cover chiuetal2026. For the modal panel regression in comparative politics and international relations, the danger is measurement error amplified by fixed effects, not negative weighting.

The Hidden Cost of Fixed Effects

Measurement Error Amplification

To see this issue, consider Polity scores. A country's Polity score is usually relatively constant over time, but there is typically a lot of cross-sectional variation. For instance, Saudi Arabia and Norway have very different scores, and most of the variable's overall signal is in those cross-country differences. Polity also includes measurement error because coders disagree, the historical record is thin, and later revisions rewrite earlier country-year values.

Country fixed effects subtract a country's average Polity score from its score in a given year, so the model uses within-country variation around its average to identify coefficients. These models are useful under time-invariant confounding. However, for a slow moving variable like Polity, the country average is not only confounding, it is also where much of the real variation is. After country fixed effects are added, the model relies on the remaining movement within countries. Some of that movement reflects real regime change. Some of it reflects coding disagreement, later revisions, or small annual changes with little substantive meaning. Fixed effects can therefore leave the researcher with a noisier version of the original variable.

The formal result follows from a standard panel model: \[ y_{it} = \alpha_i + \beta x_{it}^* + \varepsilon_{it}, \] where $y_{it}$ is the outcome for unit $i$ in period $t$, $\alpha_i$ is a unit fixed effect, $x_{it}^*$ is the true regressor, and $\varepsilon_{it}$ is the error term. The researcher observes a noisy version, \[ x_{it} = x_{it}^* + u_{it}, \] where $u_{it}$ is classical measurement error: zero mean, uncorrelated with the true regressor and the outcome error. The overall reliability of the observed regressor is \[ \lambda = \frac{\operatorname{Var}(x^*)}{\operatorname{Var}(x)}, \] the share of observed variance that is signal rather than noise.

With no fixed effects, classical measurement error attenuates the coefficient toward zero by the factor $\lambda$. With fixed effects, the quantity that matters is not the overall reliability but the reliability of the within unit variation in $x_{it}$. Call it the within reliability $\lambda_w$, and, importantly, the fixed effects estimator converges to $\beta \lambda_w$, not $\beta$. How much reliability survives the within transformation depends on how much of the regressor's variance lies between units. Let $\operatorname{ICC}^*$ be the intraclass correlation of the true regressor, \[ \operatorname{ICC}^* = \frac{\operatorname{Var}(\bar{x}_i^*)}{\operatorname{Var}(x^*)}, \] the share of variance that lies between rather than within units. A high-ICC variable barely moves within units over time, and griliches1986 show that, under classical measurement error, \[ \lambda_w = \frac{(1-\operatorname{ICC}^*)\lambda}{1-\operatorname{ICC}^*\lambda}. \] The more persistent the true regressor, the higher $\operatorname{ICC}^*$ and the lower the within reliability.

One cannot use this formula in practice because $x_{it}^*$ is unobserved and the empirical ICC is itself attenuated by measurement error. Under classical measurement error with serially uncorrelated noise and large $T$, \[ \widehat{\operatorname{ICC}} = \operatorname{ICC}^* \lambda. \] Substituting $\operatorname{ICC}^* = \widehat{\operatorname{ICC}}/\lambda$ into the Griliches-Hausman expression yields a corrected within reliability formula:

equation[equation omitted — 129 chars of source]

One needs two quantities for this computation: the empirical ICC of the observed regressor, computed directly from the data, and the overall reliability $\lambda$, from a measurement model, a second measure, published reliability evidence, or a defended range.

Equation (ref) is also a diagnostic. A slow moving regressor with a high empirical ICC can have high overall reliability, yet low within reliability. If $\widehat{\operatorname{ICC}} = 0.75$ and $\lambda = 0.90$, then \[ \lambda_w = \frac{0.90 - 0.75}{1 - 0.75} = 0.60, \] so a fixed effects model recovers only about 60 percent of the coefficient.

Figure (ref) shows the attenuation problem for five common variables in political science. Overall reliability is high for all five, so none would look problematic in a pooled regression. Once the model uses only within unit change, though, Polity and log GDP per capita lose roughly half their signal, and CINC loses nearly three quarters. Thus, a measure can be reliable in levels yet unreliable in the variation fixed effects use.

figure[figure omitted — 479 chars of source]

This derivation assumes serially uncorrelated measurement error, which leaves most of the noise in the within dimension and makes the diagnosis conservative. I take up persistent, coder- or source-specific error below.

The Scope of the Problem

Are these examples unusual? Many panel variables in political science move slowly, and high persistence is exactly when fixed effects amplify measurement error. Figure (ref) audits every numeric country-year variable in the Quality of Government and V-Dem Country-Year datasets qog2025, vdem2026, and it computes the empirical ICC\footnote{$\widehat{\operatorname{ICC}} = \operatorname{Var}(\bar{x}_i)/\operatorname{Var}(x)$} for each.

The results show that high persistence is widespread in both datasets. In the Quality of Government data, 67 percent of the variables have ICCs above 0.70 and 34 percent exceed 0.90. V-Dem is less extreme but still contains a significant percentage of highly persistent variables: 37 percent exceed 0.70 and 10 percent exceed 0.90, and many of its core institutional and regime measures sit in the “danger zone.”

figure[figure omitted — 504 chars of source]

This exploratory analysis shows the enormous potential for fixed effects attenuation in the political science literature, but not that any particular variable is severely attenuated, because severity depends on both the empirical ICC and reliability. A high-ICC variable with near-perfect reliability may retain meaningful within signal, while one with moderate reliability may not. Appendix (ref) reports estimated within-reliability for 20 commonly used variables. I draw on pemstein2018 for V-Dem indices and the relevant measurement literature otherwise.\footnote{Other sources include treier2008, johnson2013, feenstra2015, and kaufmann2011.} Nine of the 20 variables lose more than 30 percent of their coefficient under unit fixed effects, and two of them, the Corruption Perceptions Index and the trade-to-GDP ratio, lose more than 80 percent.

Two-way fixed effects make the problem worse because year fixed effects also remove common time variation, which for variables such as income or capability is often substantively meaningful. The two-way FE diagnostic replaces the unit ICC with the share of variance absorbed jointly by unit and year:

equation[equation omitted — 146 chars of source]

In a balanced panel with orthogonal unit and year components, $\widehat{\operatorname{ICC}}_{uy} = \widehat{\operatorname{ICC}}_u + \widehat{\operatorname{ICC}}_t$. In unbalanced panels, I compute the joint absorbed component directly. The difference between one-way and two-way fixed effects can be large. In the treisman2015 democracy and growth data, log GDP per capita has $\widehat{\operatorname{ICC}}_u = 0.70$ and $\widehat{\operatorname{ICC}}_t = 0.26$ with $\lambda = 0.97$. In this case, one-way country fixed effects imply $\lambda_w = 0.90$, while two-way fixed effects imply $\lambda_w^{\text{2FE}} = 0.25$, which leaves only one quarter of the coefficient recoverable under the measurement error model. These calculations identify when an estimate is likely to rely on low reliability within variation. They do not show that every high ICC variable is problematic, and the reporting standard below evaluates reliability on a case-by-case basis.

Precisely Wrong: The Inference Problem

Measurement error also changes how researchers should interpret the statistical significance of coefficients.Scholars typically assume large samples and small standard errors bring estimates closer to the “truth,” but under fixed effects the opposite can hold. The sampling distribution of the FE estimate is centered on $\beta\lambda_w$, not $\beta$, and the bias $\beta(1-\lambda_w)$ does not shrink as sample size increases. The sampling variance shrinks while the estimator remains centered on the attenuated value, so the confidence interval tightens around the wrong value. Counterintuitively, a small panel's interval may cover the true effect, while a large panel's tighter interval is significant, yet excludes it. Figure (ref) demonstrates this issue. When reliability is perfect, coverage stays near the 95 percent level as $N$ increases. But with even modest attenuation, larger samples make the problem worse. The coverage of the true value remains close to nominal for $\lambda = 0.95$ through moderate sample sizes before falling sharply, deteriorates much earlier for $\lambda = 0.90$, and collapses almost immediately for lower values of $\lambda$. Put simply, a “significant” finding may mean that one's dataset is large enough to estimate the attenuated quantity, $\beta\lambda_w$, with great precision while excluding the true effect, $\beta$. Appendix (ref) confirms this in a 243-cell simulation, where bare fixed-effects intervals cover the truth in 4.8 percent of replications.

figure[figure omitted — 467 chars of source]

The Diagnostic

Which estimate should a researcher report? The diagnostic answers this using only quantities the researcher already has: the pooled OLS coefficient, the fixed-effects coefficient, the empirical ICC of the key regressor, and, when available, a defensible reliability estimate. It proceeds in two steps. The researcher first settles the two measurement inputs, the ICC and a reliability estimate, then applies a sign test that determines which result to report.

The researcher estimates reliability from the strongest available source. For measures that report observation-level uncertainty, such as the V-Dem indices, reliability follows from the posterior uncertainty. When two independent measures of the same construct exist, as for democracy, their association supplies external information about reliability. For common variables with no attached uncertainty, published reliability estimates serve when the literature provides them. When none of these sources exists, I describe a frontier of reliability values for scholars to use.

Why not estimate the noise variance instead of bounding it? Under transitory error, contrasting estimators that difference the data at different lengths point-identifies reliability griliches1986, meijerspierdijkwansbeek2017, wansbeek2001. This strategy fails twice over for the variables in Figure (ref). Coder-based error is plausibly persistent, and persistent error attenuates short and long differences alike, so the contrast carries little information, and error as persistent as the signal carries none. This is the non-identification behind Proposition 2 below. Even transitory error leaves the contrast weakly identified, because the differences of a slow moving regressor are tiny: $\mathrm{Var}(\Delta_k x^*) = 2(1-\phi^k)\,\mathrm{Var}(x^*)$ with $\phi$ near one. Lag instruments inherit the same weakness near a unit root blundell1998. When a long panel and transitory error can be defended, Griliches-Hausman estimation simply supplies the $\lambda$ that the bounds take as input. Otherwise the frontier states the persistence assumption rather than fixing it at zero.

For the V-Dem indices I compute reliability from posterior uncertainty but report wider defended ranges, because posterior standard deviations capture coder disagreement and model uncertainty, not scale validity or systematic coder bias. The same logic answers a natural alternative: propagating the posterior draws of the index through the regression treier2008, pemstein2018. Every draw contains the measurement noise, so the fixed effects estimate in every draw is centered on $\beta\lambda_w$, and averaging over draws widens the interval around the wrong point estimate. The workflow instead uses the posterior to estimate the noise share and correct for it. Draw propagation also requires a measurement model, which most variables in Figure (ref) lack. Appendix (ref) gives the details.

For a variable with no attached uncertainty and no alternative measure, reliability is not identified, and the relevant object is the within-reliability frontier of Proposition 2, which bounds $\lambda_w$ directly under an assumption about error persistence.

The empirical ICC places the application on the attenuation scale of Equation (ref). At an ICC near 0.70 and reliability near 0.85, the corrected within-reliability falls to roughly 0.50, so fixed effects recover only about half the coefficient under the measurement-error model. The ICC tells the researcher how severe attenuation may be once reliability is specified.

With both inputs in hand, the sign test determines what kind of result to report. When the pooled OLS and fixed effects estimates share a sign and the latter is smaller in magnitude, the researcher uses the defensible range of reliability values to report an identified set. When the fixed effects estimate is larger in magnitude than pooled OLS, pooled OLS no longer bounds $|\beta|$ from above. This complement case supports the within-unit result, reported as a one-sided floor on $|\beta|$ rather than as an identified set. When the two estimates have opposite signs, the bounds do not apply, and the fixed effects estimate should be reported as a within-unit estimate.

The diagnostic changes how one interprets when coefficients shrink under fixed effects. Same-sign shrinkage does not by itself show that the pooled estimate was confounded away. It may instead show that the fixed effects estimate depends on low-reliability within-unit variation. When the resulting identified set includes both zero and substantively meaningful effects, one should not conclude that the “true” effect is zero, but rather that the available design and data cannot distinguish a null from a meaningful effect. Appendix (ref) characterizes the rule's operating characteristics.

Partial Identification

The identified set for the coefficient

The diagnostic tells the researcher when to report bounds rather than a point estimate. The two-sided identified set applies when pooled OLS and fixed effects estimates have the same sign and fixed effects is smaller in magnitude. The question is how much of the difference between the two effects reflects the fixed effects estimator removing time-invariant confounding and how much reflects measurement error attenuation. The answer is generally not point identified. When one can assume that time-invariant confounders bias pooled OLS away from zero, then pooled OLS provides an upper bound while the de-attenuated fixed effects estimate provides the lower bound. Narrow bounds mean the within-unit variation is informative even after correction, while wider bounds mean the design cannot separate attenuation from confounding. When fixed effects is larger than pooled OLS, one should instead report a one-sided floor on $|\beta|$.

proposition[Partial Identification] Suppose three conditions hold. First, pooled OLS is biased away from zero in probability limit, so that \[ |\mathrm{plim}\,\hat{\beta}_P| \geq |\beta|. \] Second, classical measurement error attenuates the fixed effects estimate toward zero, so that \[ \mathrm{plim}\,\hat{\beta}_{FE} = \beta\lambda_w, \] with $\lambda_w \in (0,1]$. Third, the true overall reliability of the regressor lies in the defended interval $[\lambda_{\min},\lambda_{\max}]$. Then a conservative outer identified set for $\beta$ is \[ \mathcal{B} = \bigcup_{\lambda \in [\lambda_{\min},\lambda_{\max}]} \left[ \min\!\left\{ \frac{\hat{\beta}_{FE}}{\lambda_w(\lambda,\widehat{\operatorname{ICC}})}, \hat{\beta}_P \right\}, \max\!\left\{ \frac{\hat{\beta}_{FE}}{\lambda_w(\lambda,\widehat{\operatorname{ICC}})}, \hat{\beta}_P \right\} \right], \] where $\lambda_w(\lambda,\widehat{\operatorname{ICC}})$ is given by Equation (ref).

Each value of $\lambda$ in the range gives one de-attenuated fixed effects coefficient; pairing it with pooled OLS and taking the union over the range produces $\mathcal{B}$. The first condition is the main assumption required. Pooled OLS bounds $|\beta|$ from above only when omitted variables bias it away from zero. In comparative politics and IR this assumption usually holds, since the leading omitted variables push the pooled estimate the same way as the hypothesized effect. However, when reverse causation or omitted variables could instead pull pooled OLS toward zero, the pooled estimate no longer bounds $|\beta|$, and the researcher reports the relaxed one-sided bound of Appendix (ref).

corollary[Sign-agreement rule] Under the conditions in Proposition (ref), pooled OLS and fixed effects share the sign of $\beta$ in probability limit. If the pooled OLS and fixed effects estimates have different signs in the sample, the bounds in Proposition (ref) do not apply.

Classical measurement error attenuates the estimate toward zero without reversing the sign. Signed confounding pushes pooled OLS farther from zero in the direction of $\beta$. Finding opposite signs in pooled and fixed effects estimates is therefore evidence that the between-unit and within-unit associations differ, and the fixed effects estimate should be reported as a within-unit estimate.

The width of $\mathcal{B}$ is itself a diagnostic. A wide identified set means that the data cannot sharply separate attenuation from confounding, so a single fixed effects coefficient creates false precision or confidence in one's conclusions.

The imbens2004 confidence interval provides inference for the identified set. It propagates sampling error in the two estimates while treating the ICC and reliability range as fixed inputs. Simulations find this plug-in conservative (Appendix (ref)).

When no measure of reliability is available

The identified set depends on the within reliability $\lambda_w$. When the researcher can defend an overall reliability $\lambda$, Equation (ref) maps that value and the empirical ICC into $\lambda_w$. But some variables provide too little information to defend a value of $\lambda$. This is common when the researcher has one observed value per unit period, no measurement model, and no second measure of the same construct. Once measurement error may persist over time, the observed data cannot separate stability in the underlying construct from stability in the error. One additional assumption can still help. If the researcher places an upper bound on how persistent the measurement error can be, the autocorrelation of the observed variable gives a lower bound on $\lambda_w$: a series more persistent than the error is allowed to be must contain some real signal.

proposition[Within reliability frontier] Let $z_{it} = x_{it} - \bar{x}_i$ be the unit demeaned observed regressor. Write \[ z_{it} = s_{it} + e_{it}, \] where $s_{it} = x^*_{it} - \bar{x}^*_i$ is the unit demeaned signal and $e_{it} = u_{it} - \bar{u}_i$ is the unit demeaned measurement error. Let \[ \lambda_w = \frac{\operatorname{Var}(s)}{\operatorname{Var}(z)}. \] Let $\rho_z(k)$, $\rho_s(k)$, and $\rho_e(k)$ be the lag $k$ autocorrelations of $z$, $s$, and $e$. Assume that the signal autocorrelation cannot exceed one, so $\rho_s(k) \leq 1$, and that the error autocorrelation is bounded above by $\psi_{\max}^{\,k}$, so $\rho_e(k) \leq \psi_{\max}^{\,k}$ for some $\psi_{\max} \in [0,1)$. Then \begin{equation} \lambda_w \geq \max\left\{0,\; \max_{k \geq 1} \frac{\rho_z(k) - \psi_{\max}^{\,k}}{1 - \psi_{\max}^{\,k}} \right\}. \end{equation} The bound is computed from the autocorrelation function of the demeaned regressor and the single assumption $\psi_{\max}$. If $\psi_{\max} = 1$, the researcher allows measurement error to be as persistent as the signal, and the useful lower bound collapses to zero. (Proof in Appendix (ref).)

The parameter $\psi_{\max}$ is the researcher's assumption about how persistent measurement error can be. Setting $\psi_{\max}=0$ assumes that measurement error does not persist over time and gives the tightest lower bound. Larger values allow more persistent error, and $\psi_{\max}=1$ leaves no useful bound. The researcher should report the frontier across a few defensible values of $\psi_{\max}$. If the lower bound stays high even when persistent error is allowed, fixed effects retain meaningful signal. If the lower bound falls near zero, the fixed effects estimate may be severely attenuated. Appendix (ref) shows that the frontier covers the true $\lambda_w$ when $\psi_{\max}$ is at least as large as the true error persistence.

corollary[Certification] Let $\rho_z(k)$ be the lag-$k$ within unit autocorrelation of the demeaned observed regressor, let \[ \underline{\lambda}_w(\psi_{\max}) = \max\!\Big\{0,\ \max_{k\ge 1}\tfrac{\rho_z(k)-\psi_{\max}^{\,k}}{1-\psi_{\max}^{\,k}}\Big\} \] be the frontier floor of Proposition (ref), and let $r \equiv |\hat\beta_{FE}/\hat\beta_P|$ be the shrinkage ratio. Suppose pooled OLS and fixed effects share a sign with $|\hat\beta_{FE}| < |\hat\beta_P|$. Then the frontier rules out the measurement error reading of the shrinkage, in the sense that the de-attenuated fixed effects estimate $\hat\beta_{FE}/\lambda_w$ cannot reach $\hat\beta_P$ for any frontier-admissible $\lambda_w$, if and only if \[ \underline{\lambda}_w(\psi_{\max}) > r . \] At $\psi_{\max}=0$ the condition is $\max_k \rho_z(k) > r$, for which $\rho_z(1) > r$ is sufficient. (Proof in Appendix (ref).)

The corollary is a simple test: compare the regressor's within unit autocorrelation to the shrinkage ratio. When the within variation is more persistent than the coefficient is shrunk, so $\rho_z(1) > r$, measurement error cannot explain the shrinkage, and the identified set lies strictly inside the cross-sectional association. The measurement error explanation requires the reverse, $\rho_z(1) \le r$: a regressor persistent in levels but transitory within units. That combination is uncommon, because slow moving regressors carry $\rho_z(1)$ near one and danger zone shrinkage makes $r$ small. The frontier therefore mainly asks whether shrinkage can be explained by measurement error, and usually shows it cannot, defending a within unit result rather than rescuing a null. The rescues this paper reports, such as cornell2020 below, come instead from a resolved reliability paired with mild shrinkage.

Applications

I work through the bureaucracy and growth study of cornell2020 to demonstrate the workflow, and then conduct a larger audit of the literature. I used Claude Opus 4.8 (Anthropic, accessed via claude.ai) in May 2026 to assist with this audit. Specifically, I used Claude to 1. identify candidate published studies whose designs fit the diagnostic's scope conditions, 2. locate and download the corresponding replication datasets, and 3. draft R code that applies the ferobust package to each application. I used the model as provided, without fine-tuning or training on my own data. I independently re-ran and verified all reported estimates, diagnostics, and verdicts against the original studies, and the full analysis is reproducible from the replication code without any AI tool. I have no competing interest in Anthropic.

Table (ref) then summarizes the audit results. The main pattern is that the within unit effect is often not identified. A single fixed effects coefficient does not reveal whether the researcher is looking at confounding removed, measurement error attenuation, a sign flip, or a complement case. The rare case in which the correction overturns a fixed effects null is credible precisely because the workflow does not rescue most null findings.

Does bureaucratic quality increase economic growth?

cornell2020 use fixed effects models to ask whether bureaucratic quality raises growth. The pooled association is large and positive, but the two-way FE estimate is small and insignificant, so they interpret the latter estimate as the more credible result. Replicating their specification, a regression of five year ahead growth on V-Dem's index of impartial administration with country and year fixed effects gives a coefficient of $0.153$ ($p = 0.158$). The estimate is positive, smaller than the pooled estimate of $0.184$ ($p = 0.002$), and indistinguishable from zero at conventional levels.

However, that fixed effects coefficient does not settle the question. It is consistent with no effect, but once measurement error is allowed, it is also consistent with an effect close to the pooled association. The sign test passes because pooled OLS and fixed effects are both positive. The issue is how much of the shrinkage reflects confounding and how much reflects attenuation. In this specification, country and year fixed effects jointly absorb $77.5$ percent of the variance in the V-Dem index, so $\widehat{\operatorname{ICC}} = 0.775$. Since V-Dem reports observation level uncertainty, I use the measurement information in the index rather than a default reliability value. At $\lambda = 0.85$, the corrected within reliability is $\hat\lambda_w = (0.85 - 0.775)/(1 - 0.775) = 0.33$. Correcting the fixed effects estimate across the defended reliability interval $[0.85, 0.95]$ gives a within country effect at or above the pooled association. The reported set is therefore the conservative one sided union $[0.184, 0.458]$ (Appendix (ref)), and the Imbens-Manski 95% interval is $[0.08, 1.01]$. Both estimates exclude zero.

Table (ref) and Figure (ref) report the results. The ordinary fixed effects interval assumes perfect measurement and includes zero. The Imbens-Manski interval uses the defended reliability range and excludes zero. The correction does not make the fixed effects estimate significant. It changes the interpretation of the fixed effects null. If the index were perfectly reliable, the null would be evidence against a within country effect. But at reliability values supported by the measurement evidence, the same estimate is consistent with an attenuated effect as large as the pooled association. Reporting only the fixed effects coefficient therefore amounts to assuming $\lambda = 1$ for an index that is not measured perfectly. The result is also sensitive to nonclassical measurement error. If measurement/coding errors in the V-Dem index is sufficiently correlated with growth, the interpretation of the case becomes ambiguous. Correlated random effects does not eliminate the problem, because its within component inherits the same attenuation as fixed effects (Appendix (ref)).

figure[figure omitted — 158 chars of source]
table[table omitted — 1,139 chars of source]

Figure (ref) depicts the workflow, and Table (ref) applies it to every application in the larger audit of the literature.

figure[figure omitted — 1,248 chars of source]
table[table omitted — 1,485 chars of source]

Table (ref) shows that the bureaucracy result is the only case in the audit where the correction overturns a fixed effects null. In the other cases, the workflow either leaves the substantive conclusion unresolved, rules out measurement error as the explanation for shrinkage, or sends the result back to fixed effects because the pooled and within unit estimates have opposite signs. To summarize, a fixed effects coefficient by itself does not tell the reader whether shrinkage reflects confounding, attenuation, or a different within unit relationship.

Scope and Sensitivity

This diagnostic framework applies to typical panel models with continuous, slow moving, mismeasured regressors. Dynamic panels raise additional issues because Nickell bias and measurement error attenuation interact, and generalized methods of moments estimators use both within and between variation arellano1991, blundell1998,nickell1981. The framework also does not carry over directly to nonlinear models, although the same mechanism remains: fixed effects can leave the coefficient identified from noisier variation hyslop2001. This literature reaches a related conclusion about this tradeoff by separating within and between slopes. The diagnostic above adds the measurement error version of that problem mundlak1978, belljones2015, kropkokubinec2020.

Also, the estimated $\lambda_w$ is a marginal diagnostic. In a multivariate regression, the coefficient is scaled by the partial within reliability of the regressor after accounting for the controls. In the application, the partial value differs from the reported $\lambda_w$ by at most about $0.04$, and measurement error in the controls is a separate issue (Appendix (ref)). The i.i.d. formula is also approximate in finite panels. Persistent measurement error raises the true within reliability, while using realized within variance can lower it. In the appendix, I conduct simulations to show that these effects partly offset, and the resulting bounds are conservative (Appendix (ref)).

When measurement error is not classical

The propositions above assume classical error. In that model, $u_{it}$ is uncorrelated with the true regressor, the outcome shock, and the outcome itself. That assumption is often unrealistic for variables created by human coders. V-Dem and Polity coders read the same historical record that researchers later use to explain outcomes, including wars, coups, economic crises, mass protests, and bureaucratic quality. If a country is downgraded for institutional quality in the same years that its economic growth deteriorates, then some of the measurement error may be correlated with the outcome. This situation would create differential measurement error rather than classical measurement error.

I treat this issue as a sensitivity problem. Let $\gamma$ be the correlation between the regressor's measurement error and the outcome. The classical model assumes $\gamma = 0$. One can construct differential measurement error bounds by varying $\gamma$ over $[-\gamma_{\max}, \gamma_{\max}]$ together with the reliability range and report the enlarged identified set (Appendix (ref)). The useful quantity is the value of $|\gamma|$ at which the substantive conclusion changes. For Cornell-Knutsen-Teorell, the “rescue” survives until about $|\gamma| = 0.15$. Beyond that point, one should interpret the case as unidentified. Such a sensitivity analysis shows how much differential error each verdict can tolerate.

Two further complications are handled in the Appendix. Serially correlated measurement error leaves more within variation as signal, so the i.i.d. bounds are conservative and tighten when persistent error is allowed (Appendix (ref)). The bounds are also robust to heterogeneous treatment effects. With classical measurement error, they cover the variance weighted estimand (Appendix (ref)). Finally, tools to conduct sensitivity analyses, such as cinelli2020 and oster2019, answer a different question. They ask how strong an omitted confounder would have to be to overturn a result. The bounds here ask how much measurement error changes what fixed effects identify.

Conclusion

Two-way fixed effects models are the implied standard in comparative politics and international relations. Under classical measurement error, however, fixed effects estimates are attenuated toward zero. For the slow-moving regressors common in this literature, that attenuation can absorb more than half of the coefficient. Substantive conclusions drawn from fixed effects shrinkage often rest on an unexamined measurement assumption: that the regressor is measured without error, or $\lambda = 1$.

The workflow in this paper unmasks this assumption. Researchers should report pooled OLS, fixed effects, the empirical ICC of the key regressor, and the within reliability that scales the attenuation. When pooled OLS and fixed effects share a sign and fixed effects is smaller in magnitude, researchers should report the identified set implied by the defended reliability range. Its width, and whether it excludes zero, should inform the substantive conclusion. When the signs differ, the bounds do not apply and the fixed effects estimate should be reported as a within-unit estimate. When the fixed effects estimate is larger, the case is a complement rather than an attenuation problem. For variables that come with no uncertainty estimates and no second measure, researchers should not invent a reliability number. If measurement error may persist over time, the data cannot separate real stability from stable error. In those cases, researchers should report the within-reliability frontier. The entire workflow is implemented in the R package ferobust. Appendix (ref) shows example output.

The recent two-way fixed effects literature documents pathologies from heterogeneous effects under binary treatments goodmanbacon2021, callaway2021, sunandabraham2021. The continuous, slow moving regressors studied here carry a different and older pathology: measurement error amplification. The bureaucracy and growth example illustrates the stakes. The fixed effects estimate is consistent with no effect, but the corrected set is also consistent with an effect as large as the cross sectional association. The fixed effects coefficient alone cannot distinguish those conclusions. Reporting the identified set, rather than a lone coefficient, lets readers see whether the within unit verdict is identified.

Data Availability Statement

The diagnostic is implemented in the open-source R package ferobust ferobust2026. It is available at \url{https://github.com/asrosenberg/ferobust} and has been submitted to CRAN.

\doublespacing \printbibliography