Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
47,891 characters · 7 sections · 11 citation commands
IMF Programs and Growth: A Source-Informed Robustness Reanalysis
\noindentHow to cite: Fern\'andez Salguero, Ricardo Alonzo. 2026. IMF Programs and Growth: A Source-Informed Robustness Reanalysis. Zenodo. DOI: \href{https://doi.org/10.5281/zenodo.21134715}{10.5281/zenodo.21134715}.
\noindentKeywords: IMF programs; economic growth; meta-analysis; meta-regression; publication bias; robust variance estimation; causal inference; within-study dependence.
The empirical literature on IMF programs and economic growth has never converged on a single answer. Some studies estimate positive growth effects, especially when programs are evaluated after stabilization, among lower-income countries, or under longer horizons. Other studies find null or negative effects, especially when the selection of countries into IMF programs is modeled explicitly or when the contractionary component of adjustment is estimated during the program period. This disagreement is expected. Countries do not enter IMF programs randomly. They enter when balance-of-payments pressures, fiscal stress, reserve losses, debt rollover problems, or political constraints make external support valuable. The counterfactual path without a program is not observed, and therefore the estimand is not a raw difference between program and non-program observations. It is a causal contrast against an unobserved trajectory.
This selection problem is central in the older evaluation literature. Goldstein and Montiel identify the methodological pitfalls of multicountry program evaluation, Bird frames the policy problem of whether programs work and under what institutional conditions, and Dicks-Mireaux, Mecagni, and Schadler evaluate IMF lending to low-income countries with attention to non-random participation goldstein1986,bird2001,dicksmireaux2000. Several influential studies then move toward explicit counterfactual or selection-aware designs. Przeworski and Vreeland estimate negative growth effects during program participation; Hardoy uses matching to reassess growth effects; Barro and Lee use political-economy instruments; Atoyan and Conway compare matching and instrumental-variable estimators; Dreher separates programs, loans, and compliance; Bas and Stone model adverse selection; Binder and Bluhm study conditional effects; Bal-Gunduz and Bird and Rowlands focus on low-income countries; and Newiak and Willems use synthetic-control evidence for non-financial programs przeworski2000,hardoy2003,barro2005,atoyan2006,dreher2006,bas2014,binder2017,balgunduz2016,bird2017,newiak2017.
Balima and Sokolova balima2021 make a major contribution by collecting 994 estimates from 36 studies and by documenting extensive heterogeneity in the IMF-growth literature. Their meta-analysis is an important benchmark because it translates a fragmented literature into a common empirical object. The question addressed here is stricter. Does the positive average reported in the broader literature survive when estimates from the same study are no longer treated as independent, when the credibility of the source design is coded from the underlying papers, and when publication bias and precision selection are evaluated jointly with heterogeneity? The answer is no. Positive estimates are present, sometimes numerous, but the positive average is fragile.
The distinction between estimates and studies is fundamental. The 994 reported estimates are not 994 independent experiments. A single article may contribute many specifications built from the same country sample, outcome definition, program measure, covariate set, and authorial research design. Repeated estimates from the same study share unobserved choices and sampling errors. Treating them as independent observations overstates precision. The stricter design in this paper therefore asks what happens when the inferential unit is the study, or when the covariance among estimates within a study is modeled rather than ignored.
A second distinction concerns identification. A method label is not a validity certificate. An instrumental-variable estimate can be weak if the instrument affects growth through channels other than IMF participation. A matching design can remain biased if selection depends on unobservables. A difference-in-differences design is credible only if the comparison path is plausible. A synthetic-control design depends on pre-fit and placebo diagnostics. A generalized evaluation estimator depends on the stability of the policy reaction function. The reanalysis therefore does not simply rank studies by whether they say “IV”, “DID”, or “PSM”. It uses a source-informed classification that assigns studies to credibility tiers while retaining a cautious interpretation of every tier.
The claim advanced here is deliberately limited. The paper does not show that IMF programs always fail, and it does not prove a precise zero effect. It shows that the existing meta-analytic evidence does not robustly establish a positive average causal effect of IMF programs on growth. In the source-informed run, effect-level random-effects specifications are often positive and statistically significant. However, the study-level estimates, multilevel CR2 estimates, RVE estimates, and precision-adjusted estimates are small and statistically indistinguishable from zero. This pattern is the empirical core of the paper.
The cleaned replication dataset contains 994 effect estimates from 36 studies. Each observation includes an effect estimate, a reported standard error, the study identifier, publication and design indicators, horizon variables, sample descriptors, and additional covariates from the original meta-analytic coding. The reanalysis adds a source-informed study classification based on the methods used in the primary studies. The classification is not a full risk-of-bias instrument. It is a structured way to avoid treating weak comparisons, partial counterfactual designs, and stronger quasi-causal designs as equivalent evidence.
The coding rule is transparent. Difference-in-differences, PSM-DID when identifiable, external or political instrumental variables with explicit selection adjustment, and synthetic-control evidence when isolated and diagnosed are treated as stronger quasi-causal designs, following the identification concerns raised in the source studies barro2005,atoyan2006,hardoy2003,newiak2017. Propensity-score matching alone, generalized evaluation estimators, generic instrumental variables with unclear exclusion restrictions, conditional-effect panel models, and panel fixed-effects designs with partial controls are treated as medium partial counterfactual designs bas2014,binder2017,balgunduz2016,bird2017. OLS, before-after comparisons, and designs that do not construct a credible counterfactual are treated as weak or without a valid counterfactual. Table (ref) gives the codebook, Table (ref) reports the distribution of the analytical sample, and Table (ref) reports the method families.
The primary statistical problem is dependence. The conventional random-effects model treats the observed estimate $\hat\theta_{ij}$ from estimate $i$ in study $j$ as \[ \hat\theta_{ij}=\theta_{ij}+e_{ij}, \qquad e_{ij}\sim N(0,s_{ij}^{2}), \] where $s_{ij}$ is the reported standard error. A simple random-effects model writes \[ \theta_{ij}=\mu+u_{ij}, \qquad u_{ij}\sim N(0,\tau^{2}). \] That specification is useful descriptively but incomplete when a study contributes many related estimates. The correlated-effects sensitivity used here allows the sampling covariance between estimates from the same study to be nonzero: \[ \operatorname{Cov}(e_{ij},e_{\ell j})=\rho s_{ij}s_{\ell j},\qquad i\neq \ell. \] The grid $\rho\in\{0,0.2,0.5,0.8\}$ is not an estimated truth. It is a sensitivity device. The value $\rho=0$ represents the conventional independence assumption, $\rho=0.2$ represents weak within-study dependence, $\rho=0.5$ represents a moderate working correlation that is plausible when specifications share samples and design choices, and $\rho=0.8$ represents a high-dependence stress test. If the positive effect only survives at low or zero correlation, it is not robust to dependence.
The implementation follows standard random-effects and multilevel meta-analysis machinery in metafor, robust variance estimation for dependent effects, and small-sample cluster-robust inference for meta-regression viechtbauer2010,hedges2010,pustejovsky2022. Publication-bias diagnostics follow the funnel-asymmetry logic of Egger, PET and PEESE meta-regression approximations, weight-function and selection approaches, and identification-aware correction for publication selection egger1997,stanley2014,vevea1995,andrewskasy2019.
The source-informed meta-regressions use the general form \[ \hat\theta_{ij}=\alpha+X_{ij}\beta+\gamma s_{ij}+u_j+v_{ij}+e_{ij}, \] where $X_{ij}$ includes causal class, method, horizon, sample, publication status, IMF-staff status, fixed-effect indicators, and design variables. The $s_{ij}$ term corresponds to PET-type publication-bias sensitivity; replacing $s_{ij}$ with $s_{ij}^{2}$ gives PEESE-type sensitivity. These regressions do not prove that one method causally changes the reported effect. They diagnose whether the literature's reported effects are systematically related to design, precision, and classification.
Outliers are handled through influence diagnostics, leave-one-study-out checks, study-level aggregation, robust-variance procedures, and mixture diagnostics rather than through mechanical deletion. This matters because extreme study means, such as very high positive averages in a small number of studies, may be substantively informative but should not dominate inference. The analysis reports the broad multiverse instead of selecting the most favorable or least favorable estimate.
The first result is that effect-level meta-analysis reproduces the positive reading of the literature. Table (ref) shows positive estimates in the full sample. The all-sample Paule-Mandel estimate is 0.257 with a confidence interval excluding zero. Yet the same table also shows why a single effect-level estimate is not enough. The estimates vary sharply across estimators and classes. In the stronger quasi-causal group, Paule-Mandel is positive, while REML is slightly negative. This estimator dependence is not a nuisance detail. It indicates that heterogeneity and weighting choices are doing substantial inferential work.
Figure (ref) shows the funnel-style distribution of effects. It is not a clean symmetric cloud around a stable mean. The empirical distribution contains many small and imprecise estimates and a wide range of positive and negative reported effects. Publication-bias diagnostics are therefore necessary, but they cannot be interpreted apart from heterogeneity.
The most important result appears at the study level. Table (ref) reports the study-level random-effects results under alternative within-study correlations. In the full sample, the point estimate declines as $\rho$ rises: 0.163 at $\rho=0$, 0.095 at $\rho=0.5$, and 0.062 at $\rho=0.8$. None of these study-level estimates is statistically significant. In the stronger quasi-causal class, the estimates are small and close to zero under moderate and high correlation. The design filters that should be most persuasive do not deliver a stable positive study-level effect.
The multilevel CR2 and RVE results reinforce the study-level pattern. Table (ref) shows that the multilevel CR2 intercept at $\rho=0.5$ is 0.036 with a confidence interval crossing zero. The RVE intercepts are about 0.043 and statistically insignificant across the working correlation grid. These estimates are among the most defensible in the package because they directly address dependent effect sizes without pretending that every reported estimate is an independent study.
Publication-bias diagnostics point in the same direction. Table (ref) reports PET and PEESE intercepts. In the full sample, the PET intercept is 0.017 and the PEESE intercept is 0.020; both are small and statistically insignificant. In the stronger quasi-causal class and in the strict DID/external-IV filter, the adjusted intercepts are slightly negative and statistically different from zero in the diagnostic regressions. These results should not be overread as evidence that IMF programs reduce growth. PET and PEESE can be unstable when heterogeneity is extreme. Their more defensible implication is that precision-adjusted estimates do not rescue the positive average.
The WAAP-like top-precision analysis is also unfavorable to the optimistic average. Table (ref) shows that the most precise quartile of estimates yields a small positive estimate in the full sample, but negative estimates in the stronger quasi-causal, medium partial counterfactual, strict DID/external-IV, and external-IV-only filters. Again, the conclusion is not a general negative causal effect. It is that the most precise estimates do not support a stable positive average.
The distinction between non-significance and equivalence is important. Table (ref) reports selected TOST results at the study level. The evidence does not generally prove equivalence under narrow bounds such as $\delta=0.10$. Equivalence appears only in some stronger-causal and higher-correlation settings under wider bounds. The correct interpretation is therefore precise: the effect is statistically indistinguishable from zero in the most defensible specifications, but exact or narrow substantive equivalence is not established uniformly.
The specification curve summarizes the entire multiverse. Table (ref) shows that many specifications are positive, and some are significantly positive. Figure (ref) shows where those estimates sit. Positive significance is concentrated in effect-level, less conservative, or less dependence-adjusted specifications. Study-level, CR2, RVE, precision-adjusted, and stronger-causal specifications move toward zero or lose significance. The specification curve therefore supports a fragility interpretation rather than a null-by-assumption interpretation.
The p-curve and caliper diagnostics in Table (ref) show that the literature is not simply a mechanical product of bunching just below five percent significance. There are many strongly significant results. But that fact does not settle the causal question. A literature can contain real statistical signal and still fail to support a stable positive causal mean once dependence, design credibility, and publication selection are accounted for.
The sample-size diagnostic in Table (ref) explains why sample size should not be used as an instrument. Log sample size is weakly related to the standard error and is strongly explained by design covariates. This violates the logic of an instrument that shifts precision without directly changing the effect-generating process. Sample size captures country coverage, sample period, method, program definition, and data quality. It belongs in diagnostic and moderator analysis, not in an exclusion restriction.
The finite-mixture and exploratory selection results add one more layer. The mixture model describes the reported effects as a distribution with several components, not as draws from a single stable positive mean. The selection-model sensitivity is negative, but it should be interpreted cautiously because parametric selection models are fragile under dependence and high heterogeneity. Together, these diagnostics support the same conclusion: the literature is heterogeneous enough that the single positive pooled mean is not the most reliable summary.
The empirical pattern is not that every specification is zero. The empirical pattern is that the positive estimate depends on how the evidence is counted. If the unit is the reported estimate, and if many related estimates from the same paper are treated as independent, the pooled effect is positive. If the unit is the study, or if within-study dependence is modeled, the effect shrinks and loses significance. If source-informed causal credibility is imposed, the strongest designs do not produce a stable positive study-level estimate. If precision-adjusted publication-bias diagnostics are applied, the positive intercept becomes small or disappears.
This interpretation is consistent with the economics of IMF programs. A program can improve external liquidity, prevent disorderly default, coordinate expectations, and anchor macroeconomic policy bird2001,dicksmireaux2000. It can also impose fiscal contraction, monetary tightening, exchange-rate adjustment, and structural reforms that depress output in the short run przeworski2000,dreher2006. The sign of the estimated growth effect should therefore depend on timing, crisis severity, program type, financing terms, conditionality, political feasibility, and the counterfactual. A short-run stabilization effect is not the same estimand as a long-run institutional or credibility effect. This is why the average treatment effect across all designs and horizons has limited substantive meaning.
The credibility classification helps interpret the pattern. PSM-only designs can be informative, but they identify effects under selection on observables. If unobserved political capacity, reserve adequacy, debt structure, or reform commitment jointly affects IMF participation and subsequent growth, PSM can remain biased. IV designs are attractive because they target endogeneity, but political or geopolitical instruments are vulnerable if they affect growth through trade, aid, diplomatic alignment, financing access, or conflict channels. DID requires credible pretrends and treatment timing; in this dataset, the DID-only evidence comes from too few studies to carry a general conclusion. GEE and strategic selection approaches are valuable but depend on stability assumptions. The aggregate literature cannot be interpreted as if all these designs identify the same clean causal estimand.
The apparent dilution of the positive effect also has a reporting interpretation. Studies that report many specifications can have disproportionate influence at the effect level. If a paper with many estimates has a positive mean, it can dominate an analysis that treats every estimate equally. Study-level inference reduces this imbalance. This is not an arbitrary conservative choice. It is a adjustment for the fact that a paper is a research design, not merely a collection of independent numbers.
Policy interpretation should therefore be modest. The evidence does not justify a broad claim that IMF programs raise growth on average. It also does not justify a broad claim that IMF programs reduce growth in all contexts. The evidence is compatible with a distribution of effects: some positive, some null, some negative. The next substantive research question is not whether the IMF works in the abstract. It is when program design, country conditions, financing constraints, and implementation capacity make growth effects more likely to be positive or negative.
The most important limitation is power. Some high-credibility filters contain few studies. DID-only evidence is too thin for a strong conclusion, and external-IV-only evidence is too heterogeneous for a precise estimate. Therefore, failure to reject zero in these subgroups should not be misread as proof of no effect. It is evidence that the current literature does not robustly establish a positive effect under stricter inferential rules.
A second limitation is unresolved heterogeneity. The high $I^{2}$ values show that the studies are not estimating a single common parameter. The reanalysis addresses this through study-level aggregation, correlated-effects sensitivity, design filters, CR2, RVE, specification curves, and publication-bias diagnostics. These are substantial improvements, but they do not fully explain heterogeneity. A more complete design would code program type, financing size, conditionality intensity, debt distress, exchange-rate regime, political institutions, pre-program crisis severity, and post-program horizon. The cleaned data contain some horizon and sample indicators, but not enough detail to estimate all mechanisms credibly.
A third limitation is classification uncertainty. The source-informed coding improves on mechanical method dummies, but it is still a structured judgment. A journal submission should include double coding by independent reviewers or a full risk-of-bias appendix. The present package provides the manual audit and study-level classification so that readers can revise the coding and rerun the analysis.
A fourth limitation concerns publication-bias adjustments. Egger, PET, PEESE, WAAP-like, trim-and-fill, and selection models each answer different questions and each can fail under high heterogeneity. They are therefore interpreted as triangulation. Their common message is that precision adjustment does not support a stable positive average; their individual point estimates should not be treated as definitive.
A fifth limitation concerns equivalence. The most defensible estimates are statistically indistinguishable from zero, but equivalence to a substantively negligible bound is not uniformly proven. The conclusion should therefore be written as a robustness claim: the positive average is not robust. It should not be written as a proof that the true effect is exactly zero.
The replication package contains the cleaned Stata dataset, the source-informed study classification, the manual source-paper audit, the R script, generated tables, generated figures, model objects, logs, and this manuscript. The script expects a data directory containing \path{data.dta}, \path{class.csv}, and \path{audit.csv}. It can be run with:
The run should report complete source-informed coverage. The package includes the output tables and session information so that the numerical results can be checked without rerunning the full pipeline.
The preferred replication sequence is simple. First, verify that all 36 studies receive a source-informed classification. Second, reproduce the effect-level random-effects models. Third, reproduce the study-level results over the $\rho$ grid. Fourth, inspect the multilevel CR2 and RVE intercepts. Fifth, examine PET, PEESE, WAAP-like, p-curve, and selection-model sensitivity as diagnostics rather than as a single decisive adjustment. Sixth, compare the specification curve to the main text to verify that the positive result is concentrated in less conservative specifications.
\begingroup
\endgroup