EconBase
← Back to paper

IMF Programs and Growth: A Source-Informed Robustness Reanalysis

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

47,891 characters · 7 sections · 11 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

IMF Programs and Growth: A Source-Informed Robustness Reanalysis

abstractThis article reassesses the meta-analytic evidence on the effect of International Monetary Fund programs on economic growth. The point of departure is the influential meta-analysis by Balima and Sokolova, which assembles 994 estimates from 36 studies and reports a positive average effect with substantial heterogeneity balima2021. The reanalysis presented here imposes four stricter requirements. It treats the study, rather than the reported estimate, as the primary inferential unit; it uses a source-informed classification of causal credibility; it models within-study dependence through study aggregation, correlated-effects sensitivity, multilevel CR2 inference, and robust variance estimation; and it evaluates publication-bias sensitivity through Egger, PET, PEESE, WAAP-like top-precision analysis, p-curve diagnostics, trim-and-fill, and exploratory selection models egger1997,stanley2014,vevea1995,andrewskasy2019. The central result is not that IMF programs reduce growth everywhere, nor that the true effect is exactly zero. The result is narrower and stronger: the positive average effect in the aggregate literature is not robust once dependence, publication selection, heterogeneity, and credibility of identification are treated as first-order concerns. In the most defensible specifications, the average effect is statistically indistinguishable from zero, while equivalence to a substantively negligible effect is only partially supported and depends on the chosen equivalence bound.

\noindentHow to cite: Fern\'andez Salguero, Ricardo Alonzo. 2026. IMF Programs and Growth: A Source-Informed Robustness Reanalysis. Zenodo. DOI: \href{https://doi.org/10.5281/zenodo.21134715}{10.5281/zenodo.21134715}.

\noindentKeywords: IMF programs; economic growth; meta-analysis; meta-regression; publication bias; robust variance estimation; causal inference; within-study dependence.

Problem

The empirical literature on IMF programs and economic growth has never converged on a single answer. Some studies estimate positive growth effects, especially when programs are evaluated after stabilization, among lower-income countries, or under longer horizons. Other studies find null or negative effects, especially when the selection of countries into IMF programs is modeled explicitly or when the contractionary component of adjustment is estimated during the program period. This disagreement is expected. Countries do not enter IMF programs randomly. They enter when balance-of-payments pressures, fiscal stress, reserve losses, debt rollover problems, or political constraints make external support valuable. The counterfactual path without a program is not observed, and therefore the estimand is not a raw difference between program and non-program observations. It is a causal contrast against an unobserved trajectory.

This selection problem is central in the older evaluation literature. Goldstein and Montiel identify the methodological pitfalls of multicountry program evaluation, Bird frames the policy problem of whether programs work and under what institutional conditions, and Dicks-Mireaux, Mecagni, and Schadler evaluate IMF lending to low-income countries with attention to non-random participation goldstein1986,bird2001,dicksmireaux2000. Several influential studies then move toward explicit counterfactual or selection-aware designs. Przeworski and Vreeland estimate negative growth effects during program participation; Hardoy uses matching to reassess growth effects; Barro and Lee use political-economy instruments; Atoyan and Conway compare matching and instrumental-variable estimators; Dreher separates programs, loans, and compliance; Bas and Stone model adverse selection; Binder and Bluhm study conditional effects; Bal-Gunduz and Bird and Rowlands focus on low-income countries; and Newiak and Willems use synthetic-control evidence for non-financial programs przeworski2000,hardoy2003,barro2005,atoyan2006,dreher2006,bas2014,binder2017,balgunduz2016,bird2017,newiak2017.

Balima and Sokolova balima2021 make a major contribution by collecting 994 estimates from 36 studies and by documenting extensive heterogeneity in the IMF-growth literature. Their meta-analysis is an important benchmark because it translates a fragmented literature into a common empirical object. The question addressed here is stricter. Does the positive average reported in the broader literature survive when estimates from the same study are no longer treated as independent, when the credibility of the source design is coded from the underlying papers, and when publication bias and precision selection are evaluated jointly with heterogeneity? The answer is no. Positive estimates are present, sometimes numerous, but the positive average is fragile.

The distinction between estimates and studies is fundamental. The 994 reported estimates are not 994 independent experiments. A single article may contribute many specifications built from the same country sample, outcome definition, program measure, covariate set, and authorial research design. Repeated estimates from the same study share unobserved choices and sampling errors. Treating them as independent observations overstates precision. The stricter design in this paper therefore asks what happens when the inferential unit is the study, or when the covariance among estimates within a study is modeled rather than ignored.

A second distinction concerns identification. A method label is not a validity certificate. An instrumental-variable estimate can be weak if the instrument affects growth through channels other than IMF participation. A matching design can remain biased if selection depends on unobservables. A difference-in-differences design is credible only if the comparison path is plausible. A synthetic-control design depends on pre-fit and placebo diagnostics. A generalized evaluation estimator depends on the stability of the policy reaction function. The reanalysis therefore does not simply rank studies by whether they say “IV”, “DID”, or “PSM”. It uses a source-informed classification that assigns studies to credibility tiers while retaining a cautious interpretation of every tier.

The claim advanced here is deliberately limited. The paper does not show that IMF programs always fail, and it does not prove a precise zero effect. It shows that the existing meta-analytic evidence does not robustly establish a positive average causal effect of IMF programs on growth. In the source-informed run, effect-level random-effects specifications are often positive and statistically significant. However, the study-level estimates, multilevel CR2 estimates, RVE estimates, and precision-adjusted estimates are small and statistically indistinguishable from zero. This pattern is the empirical core of the paper.

Identification

The cleaned replication dataset contains 994 effect estimates from 36 studies. Each observation includes an effect estimate, a reported standard error, the study identifier, publication and design indicators, horizon variables, sample descriptors, and additional covariates from the original meta-analytic coding. The reanalysis adds a source-informed study classification based on the methods used in the primary studies. The classification is not a full risk-of-bias instrument. It is a structured way to avoid treating weak comparisons, partial counterfactual designs, and stronger quasi-causal designs as equivalent evidence.

The coding rule is transparent. Difference-in-differences, PSM-DID when identifiable, external or political instrumental variables with explicit selection adjustment, and synthetic-control evidence when isolated and diagnosed are treated as stronger quasi-causal designs, following the identification concerns raised in the source studies barro2005,atoyan2006,hardoy2003,newiak2017. Propensity-score matching alone, generalized evaluation estimators, generic instrumental variables with unclear exclusion restrictions, conditional-effect panel models, and panel fixed-effects designs with partial controls are treated as medium partial counterfactual designs bas2014,binder2017,balgunduz2016,bird2017. OLS, before-after comparisons, and designs that do not construct a credible counterfactual are treated as weak or without a valid counterfactual. Table (ref) gives the codebook, Table (ref) reports the distribution of the analytical sample, and Table (ref) reports the method families.

table[table omitted — 1,241 chars of source]
table[table omitted — 495 chars of source]
table[table omitted — 641 chars of source]

The primary statistical problem is dependence. The conventional random-effects model treats the observed estimate $\hat\theta_{ij}$ from estimate $i$ in study $j$ as \[ \hat\theta_{ij}=\theta_{ij}+e_{ij}, \qquad e_{ij}\sim N(0,s_{ij}^{2}), \] where $s_{ij}$ is the reported standard error. A simple random-effects model writes \[ \theta_{ij}=\mu+u_{ij}, \qquad u_{ij}\sim N(0,\tau^{2}). \] That specification is useful descriptively but incomplete when a study contributes many related estimates. The correlated-effects sensitivity used here allows the sampling covariance between estimates from the same study to be nonzero: \[ \operatorname{Cov}(e_{ij},e_{\ell j})=\rho s_{ij}s_{\ell j},\qquad i\neq \ell. \] The grid $\rho\in\{0,0.2,0.5,0.8\}$ is not an estimated truth. It is a sensitivity device. The value $\rho=0$ represents the conventional independence assumption, $\rho=0.2$ represents weak within-study dependence, $\rho=0.5$ represents a moderate working correlation that is plausible when specifications share samples and design choices, and $\rho=0.8$ represents a high-dependence stress test. If the positive effect only survives at low or zero correlation, it is not robust to dependence.

The implementation follows standard random-effects and multilevel meta-analysis machinery in metafor, robust variance estimation for dependent effects, and small-sample cluster-robust inference for meta-regression viechtbauer2010,hedges2010,pustejovsky2022. Publication-bias diagnostics follow the funnel-asymmetry logic of Egger, PET and PEESE meta-regression approximations, weight-function and selection approaches, and identification-aware correction for publication selection egger1997,stanley2014,vevea1995,andrewskasy2019.

The source-informed meta-regressions use the general form \[ \hat\theta_{ij}=\alpha+X_{ij}\beta+\gamma s_{ij}+u_j+v_{ij}+e_{ij}, \] where $X_{ij}$ includes causal class, method, horizon, sample, publication status, IMF-staff status, fixed-effect indicators, and design variables. The $s_{ij}$ term corresponds to PET-type publication-bias sensitivity; replacing $s_{ij}$ with $s_{ij}^{2}$ gives PEESE-type sensitivity. These regressions do not prove that one method causally changes the reported effect. They diagnose whether the literature's reported effects are systematically related to design, precision, and classification.

figure[figure omitted — 1,285 chars of source]

Outliers are handled through influence diagnostics, leave-one-study-out checks, study-level aggregation, robust-variance procedures, and mixture diagnostics rather than through mechanical deletion. This matters because extreme study means, such as very high positive averages in a small number of studies, may be substantively informative but should not dominate inference. The analysis reports the broad multiverse instead of selecting the most favorable or least favorable estimate.

Evidence

The first result is that effect-level meta-analysis reproduces the positive reading of the literature. Table (ref) shows positive estimates in the full sample. The all-sample Paule-Mandel estimate is 0.257 with a confidence interval excluding zero. Yet the same table also shows why a single effect-level estimate is not enough. The estimates vary sharply across estimators and classes. In the stronger quasi-causal group, Paule-Mandel is positive, while REML is slightly negative. This estimator dependence is not a nuisance detail. It indicates that heterogeneity and weighting choices are doing substantial inferential work.

table[table omitted — 1,089 chars of source]

Figure (ref) shows the funnel-style distribution of effects. It is not a clean symmetric cloud around a stable mean. The empirical distribution contains many small and imprecise estimates and a wide range of positive and negative reported effects. Publication-bias diagnostics are therefore necessary, but they cannot be interpreted apart from heterogeneity.

figure[figure omitted — 194 chars of source]

The most important result appears at the study level. Table (ref) reports the study-level random-effects results under alternative within-study correlations. In the full sample, the point estimate declines as $\rho$ rises: 0.163 at $\rho=0$, 0.095 at $\rho=0.5$, and 0.062 at $\rho=0.8$. None of these study-level estimates is statistically significant. In the stronger quasi-causal class, the estimates are small and close to zero under moderate and high correlation. The design filters that should be most persuasive do not deliver a stable positive study-level effect.

table[table omitted — 1,795 chars of source]

The multilevel CR2 and RVE results reinforce the study-level pattern. Table (ref) shows that the multilevel CR2 intercept at $\rho=0.5$ is 0.036 with a confidence interval crossing zero. The RVE intercepts are about 0.043 and statistically insignificant across the working correlation grid. These estimates are among the most defensible in the package because they directly address dependent effect sizes without pretending that every reported estimate is an independent study.

table[table omitted — 717 chars of source]

Publication-bias diagnostics point in the same direction. Table (ref) reports PET and PEESE intercepts. In the full sample, the PET intercept is 0.017 and the PEESE intercept is 0.020; both are small and statistically insignificant. In the stronger quasi-causal class and in the strict DID/external-IV filter, the adjusted intercepts are slightly negative and statistically different from zero in the diagnostic regressions. These results should not be overread as evidence that IMF programs reduce growth. PET and PEESE can be unstable when heterogeneity is extreme. Their more defensible implication is that precision-adjusted estimates do not rescue the positive average.

table[table omitted — 1,038 chars of source]

The WAAP-like top-precision analysis is also unfavorable to the optimistic average. Table (ref) shows that the most precise quartile of estimates yields a small positive estimate in the full sample, but negative estimates in the stronger quasi-causal, medium partial counterfactual, strict DID/external-IV, and external-IV-only filters. Again, the conclusion is not a general negative causal effect. It is that the most precise estimates do not support a stable positive average.

table[table omitted — 865 chars of source]

The distinction between non-significance and equivalence is important. Table (ref) reports selected TOST results at the study level. The evidence does not generally prove equivalence under narrow bounds such as $\delta=0.10$. Equivalence appears only in some stronger-causal and higher-correlation settings under wider bounds. The correct interpretation is therefore precise: the effect is statistically indistinguishable from zero in the most defensible specifications, but exact or narrow substantive equivalence is not established uniformly.

table[table omitted — 891 chars of source]

The specification curve summarizes the entire multiverse. Table (ref) shows that many specifications are positive, and some are significantly positive. Figure (ref) shows where those estimates sit. Positive significance is concentrated in effect-level, less conservative, or less dependence-adjusted specifications. Study-level, CR2, RVE, precision-adjusted, and stronger-causal specifications move toward zero or lose significance. The specification curve therefore supports a fragility interpretation rather than a null-by-assumption interpretation.

table[table omitted — 376 chars of source]
figure[figure omitted — 210 chars of source]

The p-curve and caliper diagnostics in Table (ref) show that the literature is not simply a mechanical product of bunching just below five percent significance. There are many strongly significant results. But that fact does not settle the causal question. A literature can contain real statistical signal and still fail to support a stable positive causal mean once dependence, design credibility, and publication selection are accounted for.

table[table omitted — 294 chars of source]

The sample-size diagnostic in Table (ref) explains why sample size should not be used as an instrument. Log sample size is weakly related to the standard error and is strongly explained by design covariates. This violates the logic of an instrument that shifts precision without directly changing the effect-generating process. Sample size captures country coverage, sample period, method, program definition, and data quality. It belongs in diagnostic and moderator analysis, not in an exclusion restriction.

table[table omitted — 730 chars of source]

The finite-mixture and exploratory selection results add one more layer. The mixture model describes the reported effects as a distribution with several components, not as draws from a single stable positive mean. The selection-model sensitivity is negative, but it should be interpreted cautiously because parametric selection models are fragile under dependence and high heterogeneity. Together, these diagnostics support the same conclusion: the literature is heterogeneous enough that the single positive pooled mean is not the most reliable summary.

Interpretation

The empirical pattern is not that every specification is zero. The empirical pattern is that the positive estimate depends on how the evidence is counted. If the unit is the reported estimate, and if many related estimates from the same paper are treated as independent, the pooled effect is positive. If the unit is the study, or if within-study dependence is modeled, the effect shrinks and loses significance. If source-informed causal credibility is imposed, the strongest designs do not produce a stable positive study-level estimate. If precision-adjusted publication-bias diagnostics are applied, the positive intercept becomes small or disappears.

This interpretation is consistent with the economics of IMF programs. A program can improve external liquidity, prevent disorderly default, coordinate expectations, and anchor macroeconomic policy bird2001,dicksmireaux2000. It can also impose fiscal contraction, monetary tightening, exchange-rate adjustment, and structural reforms that depress output in the short run przeworski2000,dreher2006. The sign of the estimated growth effect should therefore depend on timing, crisis severity, program type, financing terms, conditionality, political feasibility, and the counterfactual. A short-run stabilization effect is not the same estimand as a long-run institutional or credibility effect. This is why the average treatment effect across all designs and horizons has limited substantive meaning.

The credibility classification helps interpret the pattern. PSM-only designs can be informative, but they identify effects under selection on observables. If unobserved political capacity, reserve adequacy, debt structure, or reform commitment jointly affects IMF participation and subsequent growth, PSM can remain biased. IV designs are attractive because they target endogeneity, but political or geopolitical instruments are vulnerable if they affect growth through trade, aid, diplomatic alignment, financing access, or conflict channels. DID requires credible pretrends and treatment timing; in this dataset, the DID-only evidence comes from too few studies to carry a general conclusion. GEE and strategic selection approaches are valuable but depend on stability assumptions. The aggregate literature cannot be interpreted as if all these designs identify the same clean causal estimand.

The apparent dilution of the positive effect also has a reporting interpretation. Studies that report many specifications can have disproportionate influence at the effect level. If a paper with many estimates has a positive mean, it can dominate an analysis that treats every estimate equally. Study-level inference reduces this imbalance. This is not an arbitrary conservative choice. It is a adjustment for the fact that a paper is a research design, not merely a collection of independent numbers.

Policy interpretation should therefore be modest. The evidence does not justify a broad claim that IMF programs raise growth on average. It also does not justify a broad claim that IMF programs reduce growth in all contexts. The evidence is compatible with a distribution of effects: some positive, some null, some negative. The next substantive research question is not whether the IMF works in the abstract. It is when program design, country conditions, financing constraints, and implementation capacity make growth effects more likely to be positive or negative.

figure[figure omitted — 1,142 chars of source]

Limits

The most important limitation is power. Some high-credibility filters contain few studies. DID-only evidence is too thin for a strong conclusion, and external-IV-only evidence is too heterogeneous for a precise estimate. Therefore, failure to reject zero in these subgroups should not be misread as proof of no effect. It is evidence that the current literature does not robustly establish a positive effect under stricter inferential rules.

A second limitation is unresolved heterogeneity. The high $I^{2}$ values show that the studies are not estimating a single common parameter. The reanalysis addresses this through study-level aggregation, correlated-effects sensitivity, design filters, CR2, RVE, specification curves, and publication-bias diagnostics. These are substantial improvements, but they do not fully explain heterogeneity. A more complete design would code program type, financing size, conditionality intensity, debt distress, exchange-rate regime, political institutions, pre-program crisis severity, and post-program horizon. The cleaned data contain some horizon and sample indicators, but not enough detail to estimate all mechanisms credibly.

A third limitation is classification uncertainty. The source-informed coding improves on mechanical method dummies, but it is still a structured judgment. A journal submission should include double coding by independent reviewers or a full risk-of-bias appendix. The present package provides the manual audit and study-level classification so that readers can revise the coding and rerun the analysis.

A fourth limitation concerns publication-bias adjustments. Egger, PET, PEESE, WAAP-like, trim-and-fill, and selection models each answer different questions and each can fail under high heterogeneity. They are therefore interpreted as triangulation. Their common message is that precision adjustment does not support a stable positive average; their individual point estimates should not be treated as definitive.

A fifth limitation concerns equivalence. The most defensible estimates are statistically indistinguishable from zero, but equivalence to a substantively negligible bound is not uniformly proven. The conclusion should therefore be written as a robustness claim: the positive average is not robust. It should not be written as a proof that the true effect is exactly zero.

landscape\begin{longtable}{p{0.18\linewidth}p{0.35\linewidth}p{0.39\linewidth}} \caption{Source audit.}\\ \toprule Source & Method from sources & Credibility note \\ \midrule \endfirsthead \toprule Source & Method from sources & Credibility note \\ \midrule \endhead Goldstein & Montiel (1986) & Generalized evaluation estimator foundation / multicountry evaluation pitfalls & Historically important, but assumes stable policy reaction/counterfactual; not modern causal design. \\ Hardoy (2003) & Matching and difference-in-differences matching & Better than simple PSM if DiD-matching is used; still depends on observables/common support and parallel trends. \\ Atoyan & Conway (2006) & Censored-sample, IV, and matching comparison & Important estimator comparison; results differ by estimator, so should not be collapsed into one generic quality tier. \\ Barro & Lee (2005) & IV selection using political/proximity type instruments & Causal ambition high; exclusion restriction controversial because geopolitics may affect growth and IMF access. \\ Dreher (2006) & Endogeneity-adjusted program/loan/compliance models & High relevance because it explicitly accounts for endogeneity and finds programs reduce growth; IV validity still needs audit. \\ Bas & Stone (2014) & Strategic selection model for adverse selection & Should be a separate source-informed class, not OLS/other; important for selection mechanisms. \\ Binder & Bluhm (2017) & State-dependent panel model; conditional effects and selection & Quality depends on conditional-effect specification; not equivalent to simple matching or OLS. \\ Bal-Gunduz (2016) & Propensity score matching in LIC homogeneous sample & Better PSM due to homogeneous financing events and richer selection model, but still observables-only. \\ Mumssen et al. (2013) & LIC short/longer-term impact; PSM; separates longer-term engagement and short-term financing & Useful because it separates program types/horizons, but IMF staff status should be modeled. \\ Bird & Rowlands (2017) & LIC-specific participation model with PSM & PSM-only; positive effects up to two years; credible for observables but not unobservable selection. \\ Newiak & Willems (2017) & Synthetic control for non-financial PSI programs & Should be isolated as SCM, not PSM; small number of cases and diagnostics weaken certainty. \\ Ozturk (2008/2011) & GEE applications & Not weak OLS, but relies on stable policy reaction function. \\ \bottomrule \end{longtable}

Replication

The replication package contains the cleaned Stata dataset, the source-informed study classification, the manual source-paper audit, the R script, generated tables, generated figures, model objects, logs, and this manuscript. The script expects a data directory containing \path{data.dta}, \path{class.csv}, and \path{audit.csv}. It can be run with:

quoteRscript scr/run.R --data_dir=data --out=out

The run should report complete source-informed coverage. The package includes the output tables and session information so that the numerical results can be checked without rerunning the full pipeline.

landscape\begin{longtable}{rrrlrrrr} \caption{Study classes.}\\ \toprule ID & Effects & Year & Method & Causal class & Mean & Median & Share > 0 \\ \midrule \endfirsthead \toprule ID & Effects & Year & Method & Causal class & Mean & Median & Share > 0 \\ \midrule \endhead 1 & 5 & 2006 & PSM & medium partial counterfactual & 0.017 & -0.029 & 0.400 \\ 2 & 12 & 2003 & OLS_FE_OTHER & medium partial counterfactual & 1.242 & 1.035 & 0.750 \\ 3 & 24 & 2005 & IV & stronger quasi causal & -0.079 & -0.024 & 0.042 \\ 4 & 14 & 2014 & OLS_FE_OTHER & weak no valid counterfactual & 0.049 & 0.057 & 0.857 \\ 5 & 42 & 2017 & IV & stronger quasi causal & 0.537 & 0.012 & 0.810 \\ 6 & 57 & 2017 & PSM & medium partial counterfactual & 0.825 & 1.110 & 0.789 \\ 7 & 25 & 2000 & OLS_FE_OTHER & weak no valid counterfactual & -0.185 & -0.979 & 0.440 \\ 8 & 40 & 2005 & OLS_FE_OTHER & medium partial counterfactual & -0.054 & -0.075 & 0.300 \\ 9 & 4 & 1994 & IV & medium partial counterfactual & 0.970 & 0.945 & 1.000 \\ 10 & 1 & 2000 & GEE & medium partial counterfactual & 1.374 & 1.374 & 1.000 \\ 11 & 24 & 2006 & OLS_FE_OTHER & medium partial counterfactual & -1.992 & -0.530 & 0.125 \\ 12 & 10 & 2006 & OLS_FE_OTHER & medium partial counterfactual & -0.092 & -0.012 & 0.400 \\ 13 & 8 & 2015 & OLS_FE_OTHER & weak no valid counterfactual & 2.821 & 2.136 & 1.000 \\ 14 & 3 & 1986 & BA & weak no valid counterfactual & -0.490 & -0.220 & 0.333 \\ 15 & 74 & 2016 & PSM & medium partial counterfactual & 1.372 & 1.356 & 0.986 \\ 16 & 10 & 1987 & OLS_FE_OTHER & weak no valid counterfactual & 0.370 & 0.350 & 0.800 \\ 17 & 96 & 2003 & PSM & medium partial counterfactual & 0.014 & 0.135 & 0.583 \\ 18 & 19 & 2003 & IV & medium partial counterfactual & -0.339 & -0.730 & 0.368 \\ 19 & 16 & 2003 & GEE & medium partial counterfactual & -0.608 & -0.699 & 0.125 \\ 20 & 25 & 2004 & GEE & medium partial counterfactual & 0.086 & 0.090 & 0.560 \\ 21 & 14 & 2012 & OLS_FE_OTHER & weak no valid counterfactual & 1.135 & 0.810 & 0.929 \\ 22 & 20 & 2002 & IV & stronger quasi causal & -0.137 & -0.180 & 0.300 \\ 23 & 6 & 1990 & GEE & medium partial counterfactual & -0.312 & -0.270 & 0.167 \\ 24 & 6 & 2011 & OLS_FE_OTHER & weak no valid counterfactual & -1.641 & -0.713 & 0.167 \\ 25 & 2 & 2004 & OLS_FE_OTHER & weak no valid counterfactual & 0.075 & 0.075 & 1.000 \\ 26 & 51 & 2013 & PSM & medium partial counterfactual & 2.038 & 1.760 & 1.000 \\ 27 & 2 & 2017 & PSM & medium partial counterfactual & -0.737 & -0.737 & 0.000 \\ 28 & 111 & 2016 & DID & stronger quasi causal & 0.494 & 0.357 & 0.622 \\ 29 & 12 & 2005 & OLS_FE_OTHER & weak no valid counterfactual & 1.242 & 1.035 & 0.750 \\ 30 & 80 & 2013 & OLS_FE_OTHER & weak no valid counterfactual & 0.056 & 0.004 & 0.537 \\ 31 & 2 & 1993 & IV & medium partial counterfactual & -1.329 & -1.329 & 0.000 \\ 33 & 3 & 2004 & OLS_FE_OTHER & weak no valid counterfactual & 3.066 & 2.814 & 1.000 \\ 35 & 1 & 2008 & GEE & medium partial counterfactual & -0.474 & -0.474 & 0.000 \\ 36 & 1 & 2011 & GEE & medium partial counterfactual & -2.019 & -2.019 & 0.000 \\ 37 & 54 & 2008 & IV & stronger quasi causal & 0.093 & 0.001 & 0.537 \\ 38 & 120 & 2013 & OLS_FE_OTHER & weak no valid counterfactual & 0.185 & 0.026 & 0.958 \\ \bottomrule \end{longtable}

The preferred replication sequence is simple. First, verify that all 36 studies receive a source-informed classification. Second, reproduce the effect-level random-effects models. Third, reproduce the study-level results over the $\rho$ grid. Fourth, inspect the multilevel CR2 and RVE intercepts. Fifth, examine PET, PEESE, WAAP-like, p-curve, and selection-model sensitivity as diagnostics rather than as a single decisive adjustment. Sixth, compare the specification curve to the main text to verify that the positive result is concentrated in less conservative specifications.

References

\begingroup

thebibliography{99} \bibitem{andrewskasy2019} Andrews, Isaiah, and Maximilian Kasy. 2019. “Identification of and Correction for Publication Bias.” American Economic Review 109(8): 2766--2794. \bibitem{atoyan2006} Atoyan, Ruben, and Patrick Conway. 2006. “Evaluating the Impact of IMF Programs: A Comparison of Matching and Instrumental-Variable Estimators.” The Review of International Organizations 1(2): 99--124. \bibitem{balgunduz2016} Bal-Gunduz, Yasemin. 2016. “The Economic Impact of Short-Term IMF Engagement in Low-Income Countries.” World Development 87: 30--49. \bibitem{balima2021} Balima, Hippolyte W., and Anna Sokolova. 2021. “IMF Programs and Economic Growth: A Meta-Analysis.” Journal of Development Economics 153: 102741. \bibitem{barro2005} Barro, Robert J., and Jong-Wha Lee. 2005. “IMF Programs: Who Is Chosen and What Are the Effects?” Journal of Monetary Economics 52(7): 1245--1269. \bibitem{bas2014} Bas, Muhammet A., and Randall W. Stone. 2014. “Adverse Selection and Growth under IMF Programs.” The Review of International Organizations 9(1): 1--28. \bibitem{binder2017} Binder, Michael, and Marcel Bluhm. 2017. “On the Conditional Effects of IMF Program Participation on Output Growth.” \textit{Journal of Macroeconomics} 51: 192--214. \bibitem{bird2001} Bird, Graham. 2001. “IMF Programs: Do They Work? Can They Be Made to Work Better?” \textit{World Development} 29(11): 1849--1865. \bibitem{bird2017} Bird, Graham, and Dane Rowlands. 2017. “The Effect of IMF Programmes on Economic Growth in Low Income Countries: An Empirical Analysis.” \textit{The Journal of Development Studies} 53(12): 2179--2196. \bibitem{dicksmireaux2000} Dicks-Mireaux, Louis, Mauro Mecagni, and Susan Schadler. 2000. “Evaluating the Effect of IMF Lending to Low-Income Countries.” \textit{Journal of Development Economics} 61(2): 495--526. \bibitem{dreher2006} Dreher, Axel. 2006. “IMF and Economic Growth: The Effects of Programs, Loans, and Compliance with Conditionality.” \textit{World Development} 34(5): 769--788. \bibitem{egger1997} Egger, Matthias, George Davey Smith, Martin Schneider, and Christoph Minder. 1997. “Bias in Meta-Analysis Detected by a Simple, Graphical Test.” \textit{BMJ} 315: 629--634. \bibitem{goldstein1986} Goldstein, Morris, and Peter Montiel. 1986. “Evaluating Fund Stabilization Programs with Multicountry Data: Some Methodological Pitfalls.” \textit{IMF Staff Papers} 33(2): 304--344. \bibitem{hardoy2003} Hardoy, Ines. 2003. “Effect of IMF Programmes on Growth: A Reappraisal Using the Method of Matching.” Working paper in the IMF-growth evaluation literature. \bibitem{hedges2010} Hedges, Larry V., Elizabeth Tipton, and Matthew C. Johnson. 2010. “Robust Variance Estimation in Meta-Regression with Dependent Effect Size Estimates.” \textit{Research Synthesis Methods} 1(1): 39--65. \bibitem{newiak2017} Newiak, Monique, and Tim Willems. 2017. “Evaluating the Impact of Non-Financial IMF Programs Using the Synthetic Control Method.” IMF Working Paper 17/109. \bibitem{przeworski2000} Przeworski, Adam, and James Raymond Vreeland. 2000. “The Effect of IMF Programs on Economic Growth.” \textit{Journal of Development Economics} 62(2): 385--421. \bibitem{pustejovsky2022} Pustejovsky, James E., and Elizabeth Tipton. 2022. “Meta-Analysis with Robust Variance Estimation: Expanding the Range of Working Models.” \textit{Prevention Science} 23: 425--438. \bibitem{stanley2014} Stanley, T. D., and Hristos Doucouliagos. 2014. “Meta-Regression Approximations to Reduce Publication Selection Bias.” \textit{Research Synthesis Methods} 5(1): 60--78. \bibitem{vevea1995} Vevea, Jack L., and Larry V. Hedges. 1995. “A General Linear Model for Estimating Effect Size in the Presence of Publication Bias.” \textit{Psychometrika} 60: 419--435. \bibitem{viechtbauer2010} Viechtbauer, Wolfgang. 2010. “Conducting Meta-Analyses in R with the metafor Package.” \textit{Journal of Statistical Software} 36(3): 1--48.

\endgroup