EconBase
← Back to paper

An Empirical Comparison of Weak-IV-Robust Procedures in Just-Identified Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

19,203 characters · 8 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

An Empirical Comparison of Weak-IV-Robust Procedures in Just-Identified Models

titlepage\begin{abstract} Instrumental variable (IV) regression is recognized as one of the five core methods for causal inference, as identified by Angrist-Pischke(2008). This paper compares two leading approaches to inference under weak identification for just-identified IV models: the classical Anderson--Rubin ($AR$) procedure and the recently popular \( tF \) method proposed by lee2022valid. Using replication data from the American Economic Review (AER) and Monte Carlo simulation experiments, we evaluate the two procedures in terms of statistical significance testing and confidence interval (CI) length. Empirically, we find that the $AR$ procedure typically offers higher power and yields shorter CIs than the \( tF \) method. Nonetheless, as noted by lee2022valid, \( tF \) has a theoretical advantage in terms of expected CI length. Our findings suggest that the two procedures may be viewed as complementary tools in empirical applications involving potentially weak instruments. JEL: C26, C12, C15 Keywords: Instrumental variables, Weak identification, Just-identified models, Hypothesis testing, Confidence intervals \end{abstract}

Introduction

The instrumental variable (IV) regression is one of the five most commonly used causal inference methods identified by Angrist-Pischke(2008). However, weak instruments remain persistent concerns in IV regressions across various fields. For instance, surveys by Andrews-Stock-Sun(2019), young2022, and lee2022valid find that a considerable number of IV specifications in the American Economic Review (AER) report first-stage $F$-statistics below 10.

Consider the commonly used linear IV model, with outcome $Y$, endogenous regressor $X$, and instrument $Z$:

align[align omitted — 147 chars of source]

where \( Cov(u, Z) = 0 \) and \( Cov(Z, X) \neq 0 \). The problem of conducting hypothesis tests and constructing confidence sets for the structural parameter $\beta$ with valid significance and confidence levels has been extensively studied for decades. In this context, the Anderson–Rubin ($AR$) test remains a well-established and foundational method (e.g., see anderson1949estimation, Dufour(1997), and Andrews-Stock-Sun(2019), among others). In particular, a variety of weak-identification-robust procedures have been built upon $AR$ (e.g., see Stock-Wright(2000), Kleibergen(2002), Moreira(2003), Andrews(2016), and Andrews-Mikusheva(2016)). The $AR$ test is uniformly valid under the weak instrument asymptotics of Staiger-Stock(1997) and maintains correct size regardless of the strength of the first-stage regression. Furthermore, Moreira(2009) shows that $AR$ is the uniformly most powerful unbiased test in the just-identified case.

Recently, lee2022valid propose a new weak-IV-robust method for just-identified models, termed the $tF$ procedure, which preserves the interpretability of the conventional 2SLS t-ratio while addressing its invalidity under weak identification. Building on the framework of stock2002testing, the $tF$ procedure introduces a smooth critical value (CV) function indexed by the first-stage $F$-statistic, thereby ensuring uniform size control across different levels of instrument strength. Furthermore, the authors argue that $tF$ improves upon the $AR$ procedure in certain respects—most notably by yielding confidence intervals (CIs) with shorter expected lengths than those of $AR$, whenever both are bounded.

However, empirical evidence regarding the relative performance of the two weak-IV-robust procedures remains sparse. In this paper, we conduct a performance comparison regarding the test of statistical significance and CI length between the two procedures by using both empirical dataset and simulation experiments. We find that the $AR$ procedure tends to outperform the $tF$ procedure in several key aspects.

Performance Comparison Between $AR$ and \( tF \)

Dataset for the Empirical Comparison

Our empirical analysis is based on the same AER dataset as that studied in lee2022valid, which includes all AER articles published between 2013 and 2019, excluding comments, replies, and AEA Papers and Proceedings. This yields a total of 757 articles, among which 123 contain IV regressions. Of these, 61 studies employ single-instrument (just-identified) regressions, contributing a total of 1,311 specifications of IV regressions to the dataset. Among them, 458 specifications across 39 studies have matching sample sizes in the first- and second-stage regressions. Another 19 specifications are dropped due to missing the values of $F$-statistics, resulting in a working dataset of 439 specifications from 39 studies, as used by lee2022valid.

However, we find that this dataset still contains several cases with multiple instruments or nonlinear model structures, which may violate the assumptions underlying the $AR$ and \( tF \) procedures. After removing these cases, 343 specifications from 36 studies remain. Among the remaining specifications, we are able to compute the $AR$ test for 151 specifications from 17 studies, primarily due to limitations related to data confidentiality (i.e., we exclude the studies that are based on private data). Additional details of the 17 studies are provided in the Online Appendix.

Statistical Significance

We first test the null hypothesis $H_0: \beta = 0$ for all the 151 specifications. To compare the null rejection behavior of the $AR$ and \( tF \) tests using the AER replication sample, we present Figure (ref). Let $t$ denote the conventional $t$-ratio statistic. For each specification in our dataset, the vertical axis plots the value of its standardized \( t^2 \) statistic, defined as \( \frac{t^2 / 1.96^2}{1 + t^2 / 1.96^2} \), while the horizontal axis shows the value of the corresponding standardized first-stage \( F \)-statistic, calculated as \( \frac{F / 10}{1 + F / 10} \), to facilitate full visualization of both statistics. Additionally, we plot the conventional rejection thresholds for the $t$-ratio statistics in dotted lines (that is, \( t^2 > 1.96^2 \) for the 5% significance level and \( t^2 > 2.576^2 \) for the 1% level, respectively).

In Figure (ref), black circles indicate insignificance, blue circles denote significance at the 5% level only, and red circles represent significance at the 1% level, respectively, for the $AR$ procedure. Then, we plot CV curves for the \( tF \) procedure in Figure (ref), where the solid black line represents the CV curve for the 5% level and the solid gray line represents that for the 1% level, respectively. The $tF$ CV depends on the values of both $t^2$ and $F$ by construction.

figure[figure omitted — 698 chars of source]

We observe that among the specifications deemed insignificant under the \( tF \) procedure, 47.52% are nonetheless found to be significant under the $AR$ test—corresponding to the proportion of blue and red circles located below the 5% \( tF \) CV curve (solid black line). Furthermore, among those specifications that are significant at the 5% level under the \( tF \) procedure but not at the 1% level, 50% are in fact significant at the 1% level under $AR$—represented by the proportion of red circles within the region bounded between the 5% and 1% \( tF \) CV curves (i.e., the area between the solid black and gray lines). Importantly, there are no cases in which a specification is significant under the \( tF \) procedure at a level more stringent than under $AR$, reflecting the relatively conservative nature of the \( tF \) test for testing statistical significance in the current dataset. These findings suggest that the $AR$ procedure may have a power advantage over the $tF$ procedure in a variety of empirically relevant cases.

Power Simulations

To further examine the power properties of the two procedures, we conduct simulation experiments, systematically comparing the performance of the $AR$ and \( tF \) procedures under a range of scenarios. Specifically, we plot power curves under a wide range of alternatives \( \beta - \beta_0 \), where \( \beta \) denotes the true parameter value and \( \beta_0 \) the null hypothesis. Rejection probabilities are evaluated as a function of the deviation \( \beta - \beta_0 \) across various combinations of the key nuisance parameters \( \rho \) and \( f_0 \). In particular, instrument strength is governed by \( f_0 \) via the relationship \( \mathbb{E}[F] = f_0^2 + 1 \), while the degree of endogeneity is captured by \( \rho \), defined as the correlation between the first-stage and structural regression residuals.

Further details of the simulations and the power curves are presented in the Online Appendix. We highlight several main findings below. First, the power results indicate that the $AR$ test demonstrates higher power than the \( tF \) procedure in most cases, particularly when the instrument strength is low (e.g., when $f_0$ is equal to 1, 2, or 4). Second, across all simulations, the power of the $AR$ test appears to (weakly) dominate that of the $tF$ test when the endogeneity level is relatively low (e.g., when $\rho$ is equal to -0.3, -0.1, 0.1 or 0.3). Third, the $AR$ test has a substantial power advantage for testing positive alternatives in the cases with a negative $\rho$, while having an advantage for testing negative alternatives in the cases with a positive $\rho$.

Confidence Interval Lengths

In this section, we compare CI lengths of the $AR$ and $tF$ procedures using the AER sample.

For each specification, we compute the difference in log lengths between the two procedures: \[ \ln\left( \frac{\text{length}_{tF}}{\text{length}_{AR}} \right). \] The distribution of this measure is reported in Figure (ref). Among the 151 specifications, one is excluded because \( t_{\text{AR}}^2 = 0 \), which precludes the computation of \( \hat{\rho} \) (needed for the heatmaps in Figure (ref)), and another is removed due to an invalid estimated correlation (i.e., \( |\hat{\rho}| > 1 \)). After excluding cases with missing CI length information, 127 specifications remain at the 5% significance level and 123 at the 1% level.

figure[figure omitted — 842 chars of source]

We find that the $AR$ CI is shorter than that of the $tF$ procedure in 96.85% of specifications at the 5% level and 95.93% at the 1% level. Among cases in which the $AR$ interval is longer, the average log difference is \(-5.64\) log points at the 5% level and \(-10.84\) log points at the 1% level. Conversely, when the $AR$ interval is shorter, the average log difference is 86.41 log points at the 5% level and 142.39 log points at the 1% level. Furthermore, in 48.03% of cases at the 5% level and 76.42% at the 1% level, the $tF$ interval exceeds the $AR$ interval by more than 30 log points. For context, the standard 95% CI is longer than the 90% interval by approximately \( \ln\left( \frac{1.96}{1.645} \right) \approx 0.18 \), and the 99% interval exceeds the 95% interval by approximately \( \ln\left( \frac{2.58}{1.96} \right) \approx 0.27 \), highlighting the substantive magnitude of these differences. Overall, the $AR$ procedure yields shorter CIs than the $tF$ procedure in the majority of cases. Moreover, in the current sample, when the $AR$ method outperforms, it does so substantially; when it underperforms, the extent of loss is typically modest.

Figure (ref) presents a heatmap of the log difference in CI lengths, plotted against two key diagnostic measures: the first-stage \( F \)-statistic (vertical axis) and the absolute value of the estimated residual correlation, \( |\hat{\rho}| \) (horizontal axis). To facilitate visualization, the vertical axis applies the transformation \( \frac{F/10}{1 + F/10} \), which maps the \( F \)-statistic onto the unit interval. Representative values of \( F = 104.67 \), \( 10 \), \( 2.576^2 \), and \( 1.96^2 \) correspond to transformed values of the vertical axis \( y = 0.91 \), \( 0.5 \), \( 0.4 \), and \( 0.28 \), respectively. Each point on the heatmap reflects an interpolated average of the log difference computed over neighboring observations. Red regions indicate positive values of \( \ln\left( \frac{\text{length}_{tF}}{\text{length}_{AR}} \right) \), with deeper shades signifying larger differences in favor of the $AR$ procedure. Conversely, blue regions denote negative values, where the \( tF \) procedure yields shorter intervals. Overall, the figure demonstrates that the $AR$ procedure typically delivers substantially shorter CIs than the \( tF \) procedure, particularly in areas with low first-stage \( F \)-statistics—that is, when the IV is weak.

figure[figure omitted — 335 chars of source]
figure[figure omitted — 1,050 chars of source]

Conclusion

In this paper, we conduct a performance comparison regarding the test of statistical significance and CI length between the $AR$ and $tF$ procedures by using both AER dataset and simulation experiments. We find that the $AR$ procedure tends to outperform the $tF$ procedure in several key aspects. This is also in line with the results in keane2023instrument, keane2024practical, who recommend using the $AR$ procedure rather than $t$-ratio based procedures. However, we note that as pointed out by lee2022valid, $tF$ has the theoretical advantage over $AR$ in terms of expected CI length. Therefore, we believe that the two procedures are complementary to each other. There are several potential directions for future research. First, lee2023you recently proposed a refined $tF$ procedure (the $VtF$ procedure). It is important to empirically investigate the performance of the $VtF$ procedure compared with $AR$. Second, there is a growing literature on weak-identification-robust inference under many weak instruments and non-homoskedastic errors.\footnote{E.g., see crudu2021, MS22, matsushita2024jackknife, LWZ(2023), DKM24, boot-ligtenberg(2023), N23, and lim2024dimension, among others.} As pointed out by yap2023valid, the asymptotic framework of many weak instruments is closely related to that of just-identified models. It may be, therefore, interesting to extend the empirical investigation to the applications with many weak instruments. Third, the $AR$ and $tF$ procedures studied in the paper are based on asymptotic CVs, which may not have satisfactory finite sample performance under non-homoskedastic errors. On the other hand, it is found that when implemented appropriately, bootstrap approaches may substantially improve the inference accuracy for IV models, including the cases where IVs may be rather weak.\footnote{E.g., see Moreira-Porter-Suarez(2009), Davidson-Mackinnon(2008), Davidson-Mackinnon(2010), Davidson-Mackinnon(2014b), wang2015bootstrap, Wang-Kaffo(2016), Kaffo-Wang(2017), Wang-Doko(2018), Finlay-Magnusson(2019), Roodman-Nielsen-MacKinnon-Webb(2019), mackinnon2023fast, and Wang-Zhang2024.} It may be interesting to consider the bootstrap version of the $tF$ or $VtF$ procedure for performance improvement.

\setstretch{1}

Online Appendix

\pdfbookmark[1]{Online Appendix}{appendix}