EconBase
← Back to paper

Identification of and correction for publication bias

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,289 characters · 18 sections · 94 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification of and correction for publication bias

abstractSome empirical results are more likely to be published than others. Such selective publication leads to biased estimates and distorted inference. This paper proposes two approaches for identifying the conditional probability of publication as a function of a study's results, the first based on systematic replication studies and the second based on meta-studies. For known conditional publication probabilities, we propose median-unbiased estimators and associated confidence sets that correct for selective publication. We apply our methods to recent large-scale replication studies in experimental economics and psychology, and to meta-studies of the effects of minimum wages and de-worming programs.\\ Keywords: Publication bias, replication, meta-studies,\\ identification\\ JEL Codes: C18, C12, C13

Introduction

Despite following the same protocols, replications of published experiments frequently find effects of smaller magnitude or opposite sign than those in the initial studies open2015estimating, camerer2016evaluating. One leading explanation for replication failure is publication bias Ioannidis2005, ioannidis2008most, mccrary2016conservative, ChristensenMiguel2016. Journal editors and referees may be more likely to publish results that are statistically significant, that confirm some prior belief or, conversely, that are surprising. Researchers in turn face strong incentives to select which findings to write up and submit to journals based on the likelihood of ultimate publication. Together, these forms of selectivity lead to severe bias in published estimates and confidence sets.

This paper provides, to the best of our knowledge, the first nonparametric identification results for the conditional publication probability as a function of the empirical results of a study. Once the conditional publication probability is known, we derive bias-corrected estimators and confidence sets. Finally, we apply the proposed methods to several empirical literatures.

\paragraph{Identification of publication bias} Section (ref) considers two approaches to identification. The first uses data from systematic replications of a collection of original studies, each of which applies the same experimental protocol to a new sample from the same population as the corresponding original study. Absent selectivity, the joint distribution of initial and replication estimates is symmetric. Asymmetries in this joint distribution nonparametrically identify conditional publication probabilities, assuming the latter depend only on the initial estimate.

The second approach uses data from meta-studies. Meta-studies statistically combine the estimates from multiple (published) studies to derive pooled estimates. Meta-studies are based on estimates and standard errors from these studies. Absent selectivity the distribution of estimates for high variance studies is a noisier version of the distribution for low variance studies, under an independence assumption common in the meta-studies literature. Deviations from this prediction again identify conditional publication probabilities.

Both approaches identify conditional publication probabilities up to scale. Multiplying publication probabilities by a constant factor does not change the distribution of published estimates, and likewise does not affect publication bias and size distortions.

\paragraph{Correcting for publication bias} Section (ref) discusses the consequences of selective publication for statistical inference. For selectivity known (up to scale), we propose median unbiased estimators and valid confidence sets for scalar parameters. These results allow valid inference on the parameters of each study, rather than merely on average effects across a given literature. For settings where we must estimate the degree of selectivity, we further propose Bonferroni-corrected confidence intervals which account for estimation error in the selection model. The supplement derives optimal quantile-unbiased estimators for scalar parameters of interest in the presence of nuisance parameters, as well as results on Bayesian inference.

\paragraph{Applications} Section (ref) applies the theory developed in this paper to four empirical literatures. We first use data from the experimental economics and psychology replication studies of camerer2016evaluating and open2015estimating, respectively. Estimates based on our replication approach suggest that results significant at the 5% level are over 30 times more likely to be published than are insignificant results, providing strong evidence of selectivity. Estimation based on our meta-study approach, which uses only the originally published results, yields similar conclusions.

We then consider two settings where no replication estimates are available. The first is the literature on the impact of minimum wages on employment. Estimates based on data from the meta-study by wolfson201515 suggest that results corresponding a negative significant effect of the minimum wage on employment are about 3 times more likely to be published than are insignificant results. Positive and significant effects might also be less likely to be published than negative and significant effects, but the corresponding coefficient estimates are rather noisy. Second, we consider the literature on the impact of mass deworming on child body weight. Estimates based on data from the meta-study by deworming2016 find that results appear more likely to be included in this meta-study when they do not find a significant impact of deworming, though the standard errors are large and we cannot reject the null hypothesis of no selectivity.

\paragraph{Literature} There is a large literature on publication bias; good reviews are provided by rothstein2006publication and ChristensenMiguel2016. We will discuss some of the approaches from this literature in the context of our framework below. One popular method, used in e.g. CardKrueger1995 and egger1997bias, regresses z-statistics on the inverse of the standard error and takes a non-zero intercept as evidence of publication bias. Our approach using meta-studies builds on related intuitions. Another approach in the literature considers the distribution of p-values or z-statistics across studies, and takes bunching, discontinuities, or non-monotonicity in this distribution as indication of selectivity or estimate inflation DelongLang1992, brodeur2016star. Other approaches include the “trim and fill” method duval2000nonparametric and parametric selection models iyengar1988selection, hedges1992modeling. Some precedent for our proposed corrections to inference can be found in mccrary2016conservative, while the parametric models in our applications are related to those of hedges1992modeling.

Further recent work on publication bias includes stanley2014meta, who propose to use power as a weighting criterion for meta-analyses to increase robustness to selective publication. schuemie2014interpreting suggest empirical calibration of p-values in medical research. bruns2016p and bruns2017meta discuss meta analysis in observational settings with possibly biased estimates. stanley2017finding consider non-linear meta-regressions. carter2017correcting compare different meta-analytic methods for psychological research. Recent empirical studies exploring publication bias in economics and finance include ioannidis2017power, ChenZimmermann2017, havranek2015measuring and hou2017replicating. Finally, Furukawa2017 proposes an economic model of publication bias.

\paragraph{Road map} Section (ref) introduces the setting we consider, as well as a running example. Section (ref) presents our main identification results, and discusses approaches from the literature. Section (ref) discusses bias-corrected estimators and confidence sets, assuming conditional publication probabilities are known. Section (ref) presents results for our empirical applications. All proofs are given in the supplement, which also contains details of our applications, additional empirical and theoretical results, and a stylized model of optimal publication decisions.

\paragraph{Notation} Throughout the paper, upper case letters denote random variables and lower case letters denote realizations. The latent parameter governing the distribution of observables for a given study is $\Theta$. We condition on $\Theta$ whenever frequentist objects are considered, while unconditional expectations, probabilities, and densities integrate over the population distribution of $\Theta$ across studies. Estimates are denoted by $X$, while estimates normalized by their standard deviation are denoted by $Z$. Latent studies (published or unpublished) are indexed by $i$ and marked by a superscript $*$, while published studies are indexed by $j$. Subscripts $i$ and $j$ will sometimes be omitted when clear from context.

Setting

Throughout this paper we consider variants of the following data generating process. Within an empirical literature of interest, there is a population of latent studies $i$. The true effect $\Theta_i^*$ in study $i$ is drawn from distribution $\mu$. Thus, different latent studies may estimate different true parameters. The case where all latent studies estimate the same parameter is nested by taking the distribution $\mu$ to be degenerate.

Conditional on the true effect, the result $X^*_i$ in latent study $i$ is drawn from a known continuous distribution with density $f_{X^*|\Theta^*}$. We take both $X_i^*$ and $\Theta_i^*$ to be scalar unless otherwise noted. Studies are published if $D_i=1$, which occurs with probability $p(X^*_i)$, and we observe the truncated sample of published studies (that is, we observe $X_i^*$ if and only if $D_i=1$). Publication decisions reflect both researcher and journal decisions; we do not attempt to disentangle the two. Let $I_j$ denote the index $i$ corresponding to the $j$th published study. We obtain the following model:

defi[Truncated sampling process] Consider the following data generating process for latent (unobserved) variables.\\ $(\Theta^*_i,X_i^*,D_i)$ are jointly i.i.d. across $i$, with \begin{align*} \Theta^*_i&\sim \mu\\ X^*_i | \Theta^*_i &\sim f_{X^*|\Theta^*}(x|\Theta^*_i)\\ D_i | X^*_i, \Theta^*_i &\sim Ber(p(X^*_i)) \end{align*} Let $I_0=0$, $I_j = \min \{i:\; D_i=1,\; i>I_{j-1}\}$ and $\Theta_j = \Theta_{I_j}^*$. We observe i.i.d. draws $$ X_j=X^*_{I_j}. $$

Section (ref) considers extensions of this model that allow us to identify and estimate $p(\cdot)$. Section (ref) discusses how to use knowledge of $p(\cdot)$ to perform inference on $\Theta_j$ when $X_j$ is observed. Of central importance throughout is the likelihood of observing $X_j$ given $\Theta_j$:

lem[Truncated likelihood] The truncated sampling process of Definition (ref) implies the following likelihood: \begin{equation} f_{X|\Theta}\left(x|\theta\right) = f_{X^*|\Theta^*, D}(x|\theta,1)= \frac{p\left(x\right)}{E\left[p\left(X^*_i\right)|\Theta^*_i=\theta\right]}f_{X^*|\Theta^*}\left(x|\theta\right). \end{equation}

For fixed $\theta,$ selective publication reweights the distribution of published results by $p(\cdot).$ As we consider different values of $\theta$ for fixed $x,$ by contrast, the likelihood is scaled by the publication probability for a latent study with true effect $\theta$, $E\left[p\left(X^*_i\right)|\Theta^*_i=\theta\right].$

\paragraph{Study-level covariates}

The model of Definition (ref), and in particular independence between publication decisions and $\Theta^*$ given $X^*,$ may only hold conditional on some set of observable study characteristics $W^*$. For example, journals may treat studies on particular topics, or using particular research designs, differently. Likewise, the distribution of true effects may differ across these categories. In this case Equation (ref) would have to be modified to $$f_{X|\Theta, W}\left(x|\theta, w\right) = f_{X^*|\Theta^*, W^*, D}(x|\theta,w,1)= \frac{p\left(x, w\right)}{E\left[p\left(X^*_i, W_i^*\right)|\Theta^*_i=\theta\right]}f_{X^*|\Theta^*}\left(x|\theta,w\right).$$ In our applications, for example, we consider conditioning on journal of publication and year of initial circulation of a study. For simplicity of notation, however, we suppress such additional conditioning throughout our theoretical discussion.

An illustrative example

To illustrate our setting we consider a simple example to which we will return throughout the paper. A journal receives a stream of studies $i=1,2,\ldots$ reporting experimental estimates $Z_i^*\sim N(\Theta_i^*, 1)$ of treatment effects $\Theta_i^*$, where each experiment examines a different treatment. We denote the estimates by $Z^*$ rather than $X^*$ here to emphasize that they can be interpreted as z-statistics. Denote the distribution of treatment effects across latent studies by $\mu$. Normality is in many cases a plausible asymptotic approximation; $\operatorname*{Var}(Z^*|\Theta^*)=1$ is a scale normalization. The journal publishes studies with $Z_i^*$ in the interval $\left[-1.96, 1.96 \right]$ with probability $p(Z_i^*)=.1$, while results outside this interval are published with probability $p(Z_i^*)=1$. This publication policy reflects a preference for “significant results,” where a two-sided z-test rejects the null hypothesis $\Theta^*=0$ at the 5% level. This journal is ten times more likely to publish significant results than insignificant ones. This selectivity results in publication bias: published results, whose distribution is given by Lemma (ref) above, tend to over-estimate the magnitude of the treatment effect. Published confidence intervals under-cover the true parameter value for small values of $\Theta$ and over-cover for somewhat larger values. This is demonstrated by Figure (ref), which plots the median bias, $med(\hat\Theta_j|\Theta_j=\theta)-\theta,$ of the usual estimator $\hat\Theta_j=Z_j$, as well as the coverage of the conventional 95% confidence interval $[Z_j-1.96, Z_j+1.96]$.

figure[figure omitted — 334 chars of source]

Alternative data generating processes

To clarify the implications of our model, we contrast it with two alternative data generating processes.

\paragraph{Observability} The setup of Definition (ref) assumes that we only observe the draws $X^*$ for which $D = 1$. Alternative assumptions about observability might be appropriate, however, if additional information is available. First, we might know of the existence of unpublished studies, for example from experimental preregistrations, without observing their results $X^*$. In this case, called censoring, we observe i.i.d. draws of $(Y,D)$, where $Y= D \cdot X^*$. The corresponding censored likelihood is \[ f_{Y,D|\Theta^*}(x, d|\theta^*) = d\cdot p(x) \cdot f_{X^*|\Theta^*}\left(x|\theta\right) + (1-d) \cdot (1-E[D_i| \Theta_i^*=\theta^*]). \] Second, we might additionally observe the results $X^*$ from unpublished working papers as in Francoetal2014. The likelihood in this case is \[f_{X^*,D| \Theta^*}(x,d | \theta)= p(x)^d (1-p(x))^{1-d} \cdot f_{X^*|\Theta^*}(x|\theta).\] Even under these alternative observability assumptions, the truncated likelihood ((ref)) arises as a limited information (conditional) likelihood, so identification and inference results based on this likelihood remain valid. Specifically, this likelihood conditions on publication decisions in the model with censoring, and on both publication decisions and unpublished results in the model with $X^*$ observed. Thus, while additional information about the existence or content of unpublished studies might be used to gain additional insight, the results developed below continue to apply.

\paragraph{Manipulation of results}

Our analysis assumes that the distribution of the results $X^*$ in latent studies given the true effects $\Theta^*$, $f_{X^*|\Theta^*},$ is known. This implicitly restricts the scope for researchers to inflate the results of latent studies, cf. brodeur2016star. There are, however, many forms of manipulation or “p-hacking” Simonsohnetal2014 which are accommodated by our model. In particular, if researchers conduct many independent analyses (where the results of each analysis follow known $f_{X^*|\Theta^*}$) but write up and submit only significant analyses, this is a special case of our model. More broadly, essentially any form of manipulation can be represented in a more general model where $p$ depends on both $X^*$ and $\Theta^*.$ This extension is discussed in Section (ref) below.

Identifying selection

This section proposes two approaches for identifying $p(\cdot).$ The first uses systematic replication studies. By a “replication” we mean what Clemens2015 terms a “reproduction,” obtained by applying the same experimental protocol or analysis to a new sample from the same population as the original study. For each published $X$ in a given set of studies, such replications provide an independent estimate $X^r$ governed by the same parameter $\Theta$ as the original study. Under the assumption that selectivity operates only on $X$ and not on $X^r$, we prove nonparametric identification of $p(\cdot)$ up to scale. Under the additional assumption of normally distributed estimates we also establish identification of the latent distribution $\mu$ of true effects $\Theta^*.$ The distribution $\mu$ of $\Theta^*$, and more specifically the average $E[\Theta^*]$, is the key object of interest in most meta-studies; cf. rothstein2006publication. When the studies under consideration are on the same topic, for example the effect of minimum wage increases on employment, then the average provides a natural summary of the findings of this literature.

The second approach considers meta-studies where there is variation across published studies in the standard deviation $\sigma$ of normally distributed estimates $X$ of $\Theta$, where normality can again be understood as arising from the usual asymptotic approximations. Under the assumption that the standard deviation $\sigma^*$ is independent of $\Theta^*$ in the population of latent studies, and that publication probabilities are a function of the z-statistic $Z^*=X^* / \sigma$ alone, we again show nonparametric identification of $p(\cdot)$ up to scale, as well as of $\mu$.

Identification based on systematic replication studies is considered in Section (ref). Identification based on meta-studies is considered in Section (ref). In both sections, we return to our treatment effect example to illustrate results and develop intuition. Approaches in the literature, including meta-regressions and bunching of p-values, are discussed in the context of our assumptions in Section (ref).

Systematic replication studies

We first consider the case of systematic replication studies, where both $X^*$ and $X^{*r}$ are drawn independently from the same known distribution $f_{X^*|\Theta^*}$, conditional on $\Theta^*$. In this setting the joint density $f_{X^*,X^{*r}}$, integrating out $\Theta^*$, is symmetric in its arguments. Deviations from symmetry of $f_{X,X^r}$ identify $p(\cdot)$ up to scale. We then extend this result in several ways, allowing different sample sizes for the original and replication studies as well as selection on $\Theta.$

The symmetric baseline case

We extend the model in Definition (ref) above to incorporate a conditionally independent replication draw $X^{*r}$ which is observed whenever $X^*$ is. The key implications of our model are symmetry of the joint distribution of $(X^*, X^{*r})$, and that selectivity of publication operates only on $X^*$ and not on $X^{*r}$. The latter assumption is plausible for systematic replication studies such as open2015estimating and camerer2016evaluating, but may fail in non-systematic replication settings, for instance if replication studies are published only when they “debunk” prior published results.

defi[Replication data generating process] Consider the following data generating process for latent (unobserved) variables.\\ $(\Theta^*_i, X_i^*, D_i, X_i^{*r}, )$ are jointly i.i.d. across $i$, with \begin{align*} \Theta^*_i&\sim \mu\\ X^*_i | \Theta^*_i &\sim f_{X^*|\Theta^*}(x|\Theta^*_i)\\ D_i | X^*_i, \Theta^*_i &\sim Ber(p(X^*_i))\\ X^{*r}_i | D_i, X^*_i, \Theta^*_i &\sim f_{X^*|\Theta^*}(x|\Theta^*_i). \end{align*} Let $I_0=0$, $I_j = \min \{i:\; D_i=1,\; i>I_{j-1}\}$ and $\Theta_j = \Theta_{I_j}$. We observe i.i.d. draws of $$ (X_j, X^r_j) = (X^*_{I_j}, X^{*r}_{I_j}). $$

The next result extends Lemma (ref) to derive the joint density of $(X, X^r)$.

lem[Replication Density] Consider the setup of Definition (ref). In this setup, the conditional density of $(X,X^r)$ given $\Theta$ is \begin{align*} f_{X,X^r|\Theta}(x, x^r|\theta)&= f_{X^*,X^{*r}|\Theta^*,D}(x, x^r|\theta,1)\\ &= \frac{p(x)}{E[p(X^*_i)|\Theta^*_i=\theta]}f_{X^*|\Theta^*}\left(x|\theta\right) f_{X^*|\Theta^*}\left(x^r|\theta\right). \end{align*} The marginal density of $(X,X^r)$ is $$ f_{X,X^r}(x, x^r) = \frac{p(x)}{E[p(X^*_i)]} \int f_{X^*|\Theta^*}\left(x|\theta_i^*\right) f_{X^*|\Theta^*}\left(x^r|\theta_i^*\right) d\mu(\theta^*_i). $$

This lemma immediately implies that any asymmetries in the joint distribution of $X,X^r$ must arise from the publication probability $p(\cdot).$ In particular, \[ \frac{f_{X, X^r}(b,a)}{f_{X, X^r}(a,b) } =\frac{p(b)}{p(a)},\] whenever the denominators on either side are non-zero. Using this fact, we prove that $p(\cdot)$ is nonparametrically identified up to scale.

theorem[Nonparametric identification using replication experiments] Consider the setup for replication experiments of Definition (ref), and assume that the support of $f_{X^*, X^{*r}}$ is of the form $A \times A$ for some measurable set $A$. In this setup $p(\cdot)$ is nonparametrically identified on $A$ up to scale.

\paragraph{Testable restrictions} The density derived in Lemma (ref) shows that the model of Definition (ref) implies testable restrictions. Specifically, define $h(a,b) = \log(f_{X, X^r}(b,a)) - \log(f_{X, X^r}(a,b))$. By Lemma (ref), $h(a,b) = \log(p(b)) - \log(p(a))$, and therefore \[h(a,b) + h(b,c) + h(c,a) = 0\] for any three values $a,b,c$. One could construct a nonparametric test of the model based on these restrictions and an estimate of $f_{X,X^r}$. In the applications below we opt for an alternative approach, and test restrictions on an identified model which nests the setup of Definition (ref), detailed in Section (ref) below.

\paragraph{Illustrative example (continued)} To illustrate our identification approach using replication studies, we return to the illustrative example introduced in Section (ref). In this setting, suppose that the true effect $\Theta^*$ is distributed $N(1, 1)$ across latent studies. As before, assume that $Z^*$ is $N(\Theta^*, 1)$ distributed conditional on $\Theta^*,$ that $p(Z^*)=1$ when $|Z^*| > 1.96$, and that $p(Z^*)=.1$ otherwise. Hence, results that are significantly different from zero at the 5% level based on a two-sided z-test are ten times more likely to be published than are insignificant results.

figure[figure omitted — 631 chars of source]

This setting is illustrated in Figure (ref). The left panel of this figure shows 100 random draws $(Z^*, Z^{*r})$; draws where $|Z^*|\leq 1.96$ are marked in grey, while draws where $|Z^*|> 1.96$ are marked in blue. The right panel shows the subset of draws $(Z,Z^r)$ which are published. These are the same draws as $(Z^*, Z^{*r})$, except that 90% of the draws for which $Z^*$ is statistically insignificant are deleted.

Our identification argument in this case proceeds by considering deviations from symmetry around the diagonal $Z=Z^r.$ Let us compare what happens in the regions marked $A$ and $B$. In $A$, $Z$ is statistically significant but $Z^r$ is not; in $B$ it is the other way around. By symmetry of the data generating process, the latent $(Z^*, Z^{*r})$ fall in either area with equal probability. The fact that the observed $(Z,Z^r)$ lie in region $A$ substantially more often than in region $B$ thus provides evidence of selective publication, and the exact deviation of the distribution of $(Z,Z^r)$ from symmetry identifies $p(\cdot)$ up to scale.

Generalizations and practical complications

In practice we need to modify the assumptions above to fit our applications, where the sample size for the replication often differs from that in the initial study, and the sign of the initial estimate $X$ is normalized to be positive. We thus extend our identification results to accommodate these issues.

\paragraph{Differing variances}

To account for the impact of differing sample sizes on the distribution of $X^{*r}$ relative to $X^*$, we need to be more specific about the form of these distributions. We assume that both $X^*$ and $X^{*r}$ are normally distributed unbiased estimates of the same latent parameter $\Theta^*$, and that their variances are known. The assumption of approximate normality with known variance is already implicit in the inference procedures used in most applications. Since we require normality of only the final estimate from each study, rather than the underlying data, this assumption can be justified using standard asymptotic results even in settings with non-normal data, heteroskedasticity, clustering, or other features commonly encountered in practice. Normalizing the variance of the initial estimate to one yields the following setup, where we again denote the estimate by $Z$ rather than $X$ to emphasize the normalization of the variance.

align[align omitted — 349 chars of source]

We use $\sigma$ to denote both the standard deviation as a random variable and the realized standard deviation. We again assume that results are published whenever $D_i=1$, so that \[f_{Z,Z^r,\sigma}(z, z^r, \sigma) = f_{Z^*,Z^{*r},\sigma^* | D}(z, z^r, \sigma|1).\] Allowing the replication variance $\sigma_i^*$ to differ from one takes us out of the symmetric framework of Definition (ref). Display (ref) also allows the possibility that the distribution of $\sigma_i^*$ might depend on $Z_i^*$. Dependence of $\sigma^*_i$ on $Z^*_i$ is present, for example, if power calculations are used to determine replication sample sizes, as in both open2015estimating and camerer2016evaluating. In that case, $\sigma_i^*$ is positively related to the magnitude of $Z_i^*$, but conditionally unrelated to $\Theta_i^*$.

The following corollary states that identification carries over to this setting. The proof relies on the fact that we can recover the symmetric setting by (de)convolution of $Z^r$ with normal noise, given $Z$ and $\sigma$, which then allows us to apply Theorem (ref). The assumption of normality further allows recovery of $\mu$, the distribution of $\Theta^*$.

corConsider the setup for replication experiments in display (ref). Suppose we observe i.i.d. draws of $(Z, Z^r)$. In this setup $p(\cdot)$ is nonparametrically identified on $\mathbb{R}$ up to scale, and $\mu$ is identified as well.

\paragraph{Normalized sign} A further complication is that the sign of the estimates $Z$ in our replication datasets is normalized to be positive, with the sign of $Z^r$ adjusted accordingly: see Section (ref) below for further discussion. The following corollary shows that under this sign normalization identification of $p(\cdot)$ still holds, so long as $p(\cdot)$ is symmetric.

corConsider the setup for replication experiments of display (ref). Assume additionally that $p(\cdot)$ is symmetric, $p(z)=p(-z)$, and that $f_{\sigma|Z^*}(\sigma|z)=f_{\sigma|Z^*}(\sigma|-z)$ for all $z$. Suppose that we observe i.i.d. draws of $$(W,W^r) = \operatorname*{sign}(Z) \cdot (Z, Z^r). \label{eq: W def}$$ In this setup $p(\cdot)$ is non-parametrically identified on $\mathbb{R}$ up to scale, and the distribution of $|\Theta^*|$ is identified as well.

Selection depending on $\Theta^*$ given $X^*$

Selection of an empirical result $X$ for publication might depend not only on $X$ but also on other empirical findings reported in the same manuscript, or on unreported results obtained by the researcher. If that is the case, our assumption that publication decisions are independent of true effects conditional on reported results, $D\perp \Theta^* | X^*$, may fail. Allowing for a more general selection probability $p(X^*,\Theta^*)$, we can still identify $f_{X|\Theta}$, which is the key object for bias-corrected inference as discussed in Section (ref). Consider the following setup.

align[align omitted — 373 chars of source]

Assume again that results are published whenever $D_i=1$. The assumption $D_i | Z^*_i, \Theta^*_i \sim Ber(p(Z^*_i, \Theta_i^*))$ is the key generalization relative to the setup considered before. This allows publication decisions to depend on both the reported estimate and the true effect, and allows a wide range of models for the publication process. In particular, this accommodates models where publication decisions depend on a variety of additional variables, including alternative specifications and robustness checks not reported in the replication dataset. Publication probabilities conditional on $Z^*$ and $\Theta^*$ then implicitly average over these variables, resulting in additional dependence on $\Theta^*.$ For a worked-out example of this form, see Section (ref) of the supplement.

theoremConsider the setup for replication experiments of display (ref). In this setup $f_{Z|\Theta}$ is nonparametrically identified.

The proof of Theorem (ref) implies that the joint density $f_{Z,Z^r,\sigma,\Theta}$ is identified. Under the assumptions of display (ref) the joint density of $(Z,Z^r,\sigma,\Theta)$ is \[f_{Z,Z^r,\sigma,\Theta}(z,z^r,\sigma,\theta) = \frac{p(z,\theta)}{E[p(Z^*,\Theta^*)]} \varphi(z-\theta) \tfrac{1}{\sigma}\varphi\left(\tfrac{z^r-\theta}{\sigma}\right)f_{\sigma| Z^*}(\sigma|z) \tfrac{d\mu}{d\nu}(\theta),\] where we use $\nu$ to denote a dominating measure on the support of $\Theta$. Without further restrictions $p(z,\theta)$ is not identified; we can always divide $p(z,\theta)$ by some function $g(\theta)$ and multiply $\frac{d\mu}{d\nu}(\theta)$ by the same function to get an observationally equivalent model. Theorem (ref) implies, however, that $p(z,\theta)$ is identified up to a normalization given $\theta,$ since $$\frac{f_{Z|\Theta}(z,\theta)}{f_{Z^*|\Theta^*}(z,\theta)}=\frac{p(z,\theta)}{E[p(Z^*,\Theta^*)|\Theta^*=\theta]}.$$ We can for instance impose $\sup_z p(z,\theta) =1$ for all $\theta$ to get an identified model. In our applications we consider a parametric version of this model and test $p(z,\theta)=p(z)$ as a specification check on our baseline model.

Meta-studies

We next consider identification using meta-studies. Suppose that studies report normally distributed estimates $X^*$ with mean $\Theta^*$ and standard deviation $\sigma^{*}$, and that selectivity of publication is based on the z-statistic $Z^*=X^*/\sigma^*$. The key identifying assumption is that $\Theta^*$ is statistically independent of $\sigma^{*}$ across studies, so studies with larger sample sizes do not have systematically different estimands. Under this assumption, the distribution of $X^*$ conditional on a larger value $\sigma^{*}=\sigma_1$ is equal to the convolution of normal noise of variance $\sigma_1^2-\sigma_2^2$ with the distribution of $X^*$ conditional on a smaller value $\sigma^{*}=\sigma_2$. Deviations from this equality for the observed distribution $f_{X|\sigma}$ identify $p(\cdot)$ up to scale.

defi[Meta-study data generating process] Consider the following data generating process for latent (unobserved) variables.\\ $(\sigma_i^*, \Theta^*_i,X_i^*,D_i)$ are jointly i.i.d. across $i$, such that \begin{align*} \sigma_i^* &\sim \mu_\sigma\\ \Theta_i^* | \sigma_i^* & \sim \mu_\Theta\\ X_i^* | \Theta_i^*, \sigma_i^* &\sim N(\Theta_i^*, \sigma_i^{*2})\\ D_i | X^*_i, \Theta_i^*, \sigma_i^* &\sim Ber(p(X^*_i/\sigma_i^*)) \end{align*} Let $I_0=0$, $I_j = \min \{i:\; D_i=1,\; i>I_{j-1}\}$ and $\Theta_j = \Theta_{I_j}$. We observe i.i.d. draws of $$ (X_j, \sigma_j) = (X^*_{I_j}, \sigma^*_{I_j}). $$ Define $Z_i^*=\frac{X_i^*}{\sigma_i^*}$ and $Z_j=\frac{X_j}{\sigma_j}$.

A key object for identification of $p(\cdot)$ in this setting is the conditional density $f_{Z|\sigma}$.

lem[Meta-study density] Consider the setup of definition (ref). The conditional density of $Z$ given $\sigma$ is \[f_{Z|\sigma}(z|\sigma) = \frac{p(z)}{E[p(Z^*)|\sigma]} \int \varphi(z-\theta / \sigma) d\mu(\theta). \]

We build on Lemma (ref) to prove our main identification result for the meta-studies setting. Lemma (ref) implies that, for $\sigma_2>\sigma_1$, \[ \frac{f_{Z|\sigma}(z|\sigma_2)}{f_{Z|\sigma}(z|\sigma_1)} =\frac{E[p(Z^*)|\sigma=\sigma_1]}{E[p(Z^*)|\sigma=\sigma_2]} \cdot \frac{\int \varphi(z-\theta / \sigma_2) d\mu(\theta)}{\int \varphi(z-\theta / \sigma_1) d\mu(\theta)}, \] where the first term on the right hand side does not depend on $z$. Since $f_{Z|\sigma}(z|\sigma_2) / f_{Z|\sigma}(z|\sigma_1)$ is identified, this suggests we might be able to invert this equality to recover $\mu$, which would then immediately allow us to identify $p(\cdot)$. The proof of Theorem (ref) builds on this idea, considering $\partial_\sigma \log(f_{Z|\sigma}(z|\sigma))$.

theorem[Nonparametric identification using meta-studies] Consider the setup for experiments with independent variation in $\sigma$, described by Definition (ref). Suppose that the support of $\sigma$ contains an open interval. Then $p(\cdot)$ is identified up to scale, and $\mu$ is identified as well.

\paragraph{Illustrative example (continued)}

As before, assume that $\Theta^*$ is $N(1, 1)$ distributed. Suppose further that $\sigma^*$ is independent of $\Theta^*$ across latent studies, and that $X^*$ is $N(\Theta^*, \sigma^*)$ distributed conditional on $\Theta^*,$ $\sigma^*$. Let $p(X^*/\sigma^*)=1$ when $|X^*/\sigma^*| > 1.96$, $p(X^*/\sigma^*)=.1$ otherwise. Thus, results which differ significantly from zero at the 5% level are again ten times as likely to be published as insignificant results. This setting is illustrated in Figure (ref). The left panel of this figure shows 100 random draws $(X^*, \sigma^*)$; draws where $|X^*/\sigma^*|\leq 1.96$ are marked in grey, while draws where $|X^*/\sigma^*|> 1.96$ are marked in blue. The right panel shows the subset of draws $(X,\sigma)$ which are published, where 90% of statistically insignificant draws are deleted.

Compare what happens for two different values of the standard deviation $\sigma$, marked by $A$ and $B$ in Figure (ref). By the independence of $\sigma^*$ and $\Theta^*$, the distribution of $X^*$ for larger values of $\sigma^*$ is a noised up version of the distribution for smaller values of $\sigma^*$. To the extent that the same does not hold for the distribution of published $X$ given $\sigma$, this must be due to selectivity in the publication process. In this example, statistically insignificant observations are “missing” for larger values $\sigma.$ Since publication is more likely when $|X^*/\sigma^*| > 1.96$, the estimated values $X$ tend to be larger on average for larger values of $\sigma$, and the details of how the conditional distribution of $X$ given $\sigma$ varies with $\sigma$ will again allow us to identify $p(\cdot)$ up to scale.

figure[figure omitted — 569 chars of source]

\paragraph{Normalized sign} In some of our applications the sign of the reported estimates $X$ is again normalized to be positive. The following corollary shows that $p(\cdot)$ remains identified under this sign normalization provided it is symmetric in its argument.

corConsider the setup of Definition (ref). Assume additionally that $p(\cdot)$ is symmetric, i.e., $p(x/\sigma)=p(-x/\sigma)$. Suppose that we observe i.i.d. draws of $(|X|, \sigma)$. In this setup $p(\cdot)$ is nonparametrically identified on $\mathbb{R}$ up to scale, and the distribution of $|\Theta^*|$ is identified as well.

\paragraph{Dependence on $\sigma^*$} Publication decisions might depend not only on the z-statistic $Z^*$, but also on the standard deviation $\sigma^*$. Consider the setup of Definition (ref) modified such that $$D_i | X^*_i, \Theta_i^*, \sigma_i^* \sim Ber(p(X^*_i/\sigma_i^*) \cdot q(\sigma_i^*)).$$ Theorem (ref) immediately implies identification of the function $p(\cdot)$ for this generalized setup. The generalized setup is observationally equivalent to the model of Definition (ref) with the distribution of $\sigma^*$ reweighted by $q(\cdot)/E[q(\sigma^*)]$.

Relation to approaches in the literature

Various approaches to detect selectivity and publication bias have been proposed in the literature. We briefly analyze some of the these approaches in our framework. First, we discuss to what extent we should expect the results of significance tests to “replicate” in a sense considered in the literature, and show that the probability of such replication may be low even in the absence of publication bias. Second, we discuss meta-regressions, and show that while they provide a valid test of the null of no selectivity under our meta-study assumptions, they are difficult to interpret under the alternative. Third, we consider approaches based on the distribution of p-values or z-statistics, and analyze the extent to which bunching or discontinuities of this distribution provide evidence for selectivity or inflation of estimates.

\paragraph{Should results “replicate?”} The findings of recent systematic replication studies such as open2015estimating and camerer2016evaluating are sometimes interpreted as indicating an inability to “replicate the results” of published research. In this setting, a “result” is understood to “replicate” if both the original study and its replication find a statistically significant effect in the same direction. The share of results which replicate in this sense is prominently discussed in camerer2016evaluating. Our framework suggests, however, that the probability of replication in this sense might be low even without selective publication or other sources of bias.

Consider the setup for replication experiments in display (ref) with constant publication probability $p(\cdot)$, so that publication is not selective and $f_{Z,Z^r} = f_{Z^*, Z^{r*}}$. For illustration, assume further that $\sigma^{*}\equiv 1$. The probability that a result replicates in the sense described above is

multline*[multline* omitted — 343 chars of source]

If the true effect is zero in all studies then this probability is $0.025.$ If the true effect in all studies is instead large, so that $|\Theta^*| > M$ with probability one for some large $M$, then the probability of replication is approximately one. Thus, the probability that results replicate in this sense gives little indication of whether selective publication or some other source of bias for published research is present unless we either restrict the distribution of true effects or observe replication frequencies less than 0.025. Strengths and weaknesses of alternative measures of replication are discussed in Simonsohn2015, Gilbertetal2016, and Patiletal2016.

\paragraph{Meta-regressions}

A popular test for publication bias in meta-studies CardKrueger1995, egger1997bias uses regressions of either of the following forms:

align*[align* omitted — 163 chars of source]

where we use $E^*$ to denote best linear predictors. The following lemma is immediate.

lemUnder the assumptions of Definition (ref), if $p(\cdot)$ is constant then \begin{align*} E^*[X | 1, \sigma] = E[\Theta^*], E^*\left [Z | 1, \tfrac{1}{\sigma} \right ] = E[\Theta^*]\cdot \tfrac{1}{\sigma} \end{align*}

As this lemma confirms, meta-regressions can be used to construct tests for the null of no publication bias. In particular, absent publication bias $\beta_0 = 0$ and $\gamma_1=0$, so tests for these null hypotheses allow us to test the hypothesis of no publication bias, though there are some forms of selectivity against which such tests have no power. As also noted in the previous literature, absent publication bias the coefficients $\beta_1$ and $\gamma_0$ recover the average of $\Theta^*$ in the population of latent studies. While these coefficients are sometimes interpreted as selection-corrected estimates of the mean effect across studies DoucouliagosStanley2009, ChristensenMiguel2016, this interpretation is potentially misleading in the presence of publication bias. In particular, the conditional expectation $E[X | 1, \sigma]$ is nonlinear in both $\sigma$ and $1/\sigma$, which implies that $\beta_0$, $\gamma_1$ are generally biased as estimates of $E[\Theta^*]$.\footnote{Stanley2008 and DoucouliagosStanley2009 note this bias but suggest that one can still use $H_0:\gamma_1=0$ to test the hypothesis of zero true effect if there is no heterogeneity in the true effect $\Theta^*$ across latent studies.} To illustrate the resulting complications, we discuss a simple example with one-sided significance testing in Section (ref) of the supplement.\footnote{A further complication is that meta-regression coefficients are not interpretable in settings with sign-normalized estimates, as in two of our applications. See Section (ref) of the supplement for further discussion.}

\paragraph{The distribution of p-values and z-statistics} Another approach in the literature considers the distribution of p-values, or the corresponding z-statistics, across published studies. For example, Simonsohnetal2014 analyze whether the distribution of p-values in a given literature is right- or left-skewed. brodeur2016star compile 50,000 test results from all papers published in the American Economic Review, the Quarterly Journal of Economics, and the Journal of Political Economy between 2005 and 2011, and analyze their distribution to draw conclusions about distortions in the research process.

Under our model, absent selectivity of the publication process the distribution $f_Z$ is equal to $f_{Z^*}$. If we additionally assume that $Z^*|\Theta^* \sim N(\Theta^*,1)$ and $\Theta^* \sim \mu$, this implies that $$ f_Z(z) = f_{Z^*}(z) = (\pi \ast \varphi)(z) = \int \varphi(z-\theta) d\mu(\theta). $$ This model has testable implications, and requires that the deconvolution of $f_Z$ with a standard normal density $\varphi$ yield a probability measure $\mu$. This implies that the density $f_{Z^*}$ is infinitely differentiable. If selectivity is present, by contrast, then \[f_Z(z) = \frac{p(z)}{E[p(Z^*)]} \cdot f_{Z^*}(z), \] and any discontinuity of $f_Z(z)$ (for instance at critical values such as $z=1.96$) identifies a corresponding discontinuity of $p(z)$ and indicates the presence of selectivity: $$\frac{\lim_{z \downarrow z_0} f_Z(z)}{\lim_{z \uparrow z_0} f_Z(z)}= \frac{\lim_{z \downarrow z_0} p(z)}{\lim_{z \uparrow z_0} p(z)}.$$ If we impose that $p(\cdot)$ is a step function, for example, then this argument allows us to identify $p(\cdot)$ up to scale.

The density $f_{Z^*}$ also precludes excessive bunching, since for all $k\geq 0$ and all $z,$ $\partial_z^k f_{Z^*}(z) \leq \sup_z \partial_z^k \varphi(z)$ and $\partial_z^k f_{Z^*}(z) \geq \inf_z \partial_z^k \varphi(z)$ so that in particular $f_{Z^*}(z) \leq \varphi(0)$ and $f''_{Z^*}(z) \geq \varphi''(0)=-\varphi(0)$ for all $z$. Spikes in the distribution of $Z$ thus likewise indicate the presence of selectivity or inflation.

Unlike our model, which focuses on selection, brodeur2016star are interested in potential inflation of test results by researchers, and in particular in non-monotonicities of $f_Z$ which cannot be explained by monotone publication probabilities $p(z)$ alone. They construct tests for such non-monotonicities based on parametrically estimated distributions $f_{Z^*}$.

Corrected inference

This section derives median unbiased estimators and valid confidence sets for scalar parameters $\theta.$ For most of the section we assume $p(\cdot)$ is known; corrections accounting for estimation error in $p(\cdot)$ are discussed at the end of the section. As in our identification results, $f_{X^*|\Theta^*}$ is assumed known throughout. The supplement extends these results to derive optimal estimators for scalar components of vector-valued $\theta,$ and analyzes Bayesian inference under selective publication. While our identification results in the last section relied on an empirical Bayes perspective, which assumed that $\Theta_i^*$ was drawn from some distribution $\mu,$ this section considers standard frequentist results which hold conditional on $\Theta$. This reflects the different question at hand: Estimability of $p(\cdot)$, as considered in Section (ref), requires multiple observations of studies $j$ with potentially heterogeneous estimands $\Theta_j$. In this section, by contrast, we are interested in valid inference on $\Theta$ for a given study $j$, and so condition on $\Theta_j$. Conditioning on $\Theta_j$ corresponds to standard notions of bias and size control.

Selective publication reweights the distribution of $X$ by $p(\cdot).$ To obtain valid estimators and confidence sets, we need to correct for this reweighting. To define these corrections denote the cdf for published results $X$ given true effect $\Theta$ by $F_{X|\Theta}.$ For $f_{X|\Theta}$ the density of published results derived in Lemma (ref), $$ F_{X|\Theta}(x|\theta)= \int_{-\infty}^x f_{X|\Theta}(\tilde x|\theta) d\tilde x=\frac{1}{E[p(X^*) | \Theta^*=\theta]} \int_{-\infty}^x p(\tilde x)f_{X^*|\Theta^*}(\tilde x|\theta)d\tilde x. $$

For many distributions $f_{X^*|\Theta^*},$ and in particular in the leading normal case (see Lemma (ref) below) this cdf is strictly decreasing in $\theta$. Using this fact we can adapt an approach previously applied by, among others, D. Andrews1993 and StockWatson1998 and invert the cdf as a function of $\theta$ to construct a quantile-unbiased estimator. In particular, if we define $\hat{\theta}_{\alpha}\left(x\right)$ as the solution to

equation[equation omitted — 122 chars of source]

so $x$ lies at the $\alpha$-quantile of the distribution implied by $\hat{\theta}_{\alpha}\left(x\right)$, then $\hat{\theta}_{\alpha}\left(X\right)$ is an $\alpha$-quantile unbiased estimator for $\theta$.

theoremIf for all $x$, $F_{X|\Theta}(x|\theta)$ is continuous and strictly decreasing in $\theta,$ tends to one as $\theta\to-\infty$, and tends to zero as $\theta\to\infty,$ then $\hat{\theta}_{\alpha}(x)$ as defined in ((ref)) exists, is unique, and is continuous and strictly increasing for all $x.$ If, further, $F_{X|\Theta}(x|\theta)$ is continuous in $x$ for all $\theta$ then $\hat{\theta}_{\alpha}(X)$ is $\alpha$-quantile unbiased for $\theta$ under the truncated sampling setup of Definition (ref), \[ P \left( \hat{\theta}_{\alpha}\left(X\right)\le\theta | \Theta=\theta \right) =\alpha~\text{for all}~\theta. \]

If $f_{X^{*}|\Theta^{*}}\left(x|\theta\right)$ is normal, as in our applications, then the assumptions of Theorem (ref) hold whenever $p(x)$ is strictly positive for all $x$ and almost everywhere continuous.

lemIf the distribution of latent draws $X^*$ conditional on $(\Theta^*,\sigma^*)$ is $N(\Theta^*,\sigma^{*2})$, $p(x)>0$ for all $x$, and $p(\cdot)$ is almost everywhere continuous, then the assumptions of Theorem (ref) are satisfied.

These results allow straightforward frequentist inference that corrects for selective publication. In particular, using Theorem (ref) we can consider the median-unbiased estimator $\hat{\theta}_{\frac{1}{2}}\left(X\right)$ for $\theta$, as well as the equal-tailed level $1-\alpha$ confidence interval

equation[equation omitted — 136 chars of source]

This estimator and confidence set fully correct the bias and coverage distortions induced by selective publication. Other selection-corrected confidence intervals are also possible in this setting. For example, provided the density $f_{X^*|\Theta^*}(x|\theta)$ belongs to an exponential family one can form confidence intervals by inverting uniformly most powerful unbiased tests as in fithian2014optimal. Likewise, one can consider alternative estimators, such as the weighted average risk-minimizing unbiased estimators considered in muellerwang2015, or the MLE based on the truncated likelihood $f_{X|\Theta}.$

\paragraph{Illustrative example (continued)}

To illustrate these results, we return to the treatment effect example discussed above. Figure (ref) plots the median unbiased estimator, as well as upper and lower 95% confidence bounds as a function of $X$ for the same publication probability $p(\cdot)$ considered above. We see that the median unbiased estimator lies below the usual estimator $\hat{\theta}=X$ for small positive $X$ but that the difference is eventually decreasing in $X$. The truncation-corrected confidence interval shown in Figure (ref) has exactly correct coverage, is smaller than the usual interval for small $X$, wider for moderate values $X$, and essentially the same for $X\ge5$.

figure[figure omitted — 504 chars of source]

We do not recommend adjusting publication standards to reflect these corrections. If publication probabilities in this example were based on more stringent critical values, for instance, then the corrections discussed above would need to be adjusted. Instead, the purpose of these corrections is to allow readers of published research to draw valid inferences, taking the publication rule as given. The publication rule itself can then be chosen on other grounds, for example to maximize social welfare or provide incentives to researchers. We briefly discuss the question of optimal publication rules in the conclusion, as well as in Section (ref) supplement.

In this example, our approach is closely related to the correction for selective publication proposed by mccrary2016conservative. There, the authors propose conservative tests derived under an extreme form of publication bias in which insignificant results are never published. If we consider testing the null hypothesis that $\theta$ is equal to zero, and calculate our equal-tailed confidence interval under the publication probability $p(\cdot)$ implied by the model of mccrary2016conservative, then our confidence interval contains zero if and only if the test of mccrary2016conservative fails to reject.

\paragraph{Estimation Error in $p(\cdot)$} Thus far our corrections have assumed the conditional publication probability is known. If $p\left(\cdot\right)$ is instead estimated with error, median unbiased estimation is challenging but constructing valid confidence sets for $\theta$ is straightforward.

Suppose we parameterize the conditional publication probability by a finite dimensional parameter $\beta$, and let $\hat{\theta}_{\alpha}\left(X_{i};\beta\right)$ be the $\alpha$-quantile unbiased estimator under $\beta$. For many specifications of $p\left(\cdot\right)$, and in particular for those used in our applications below, $\hat{\theta}_{\alpha}\left(x;\beta\right)$ is continuously differentiable in $\beta$ for all $x$. If we have a consistent and asymptotically normal estimator $\hat{\beta}$ for $\beta$, for $0<\delta<\alpha$, consider the interval \[ \left[\hat{\theta}_{\frac{\alpha-\delta}{2}}\left(X;\hat{\beta}\right)-c_{1-\frac{\delta}{2}}\hat{\sigma}_{L}\left(X\right),\hat{\theta}_{1-\frac{\alpha-\delta}{2}}\left(X;\hat{\beta}\right)+c_{1-\frac{\delta}{2}}\hat{\sigma}_{U}\left(X\right)\right] \] where $c_{1-\frac{\delta}{2}}$ is the level $1-\frac{\delta}{2}$ quantile of the standard normal distribution while $\hat{\sigma}_{L}\left(x\right)$ and $\hat{\sigma}_{U}\left(x\right)$ are delta-method standard errors for $\hat{\theta}_{\frac{\alpha-\delta}{2}}\left(x;\hat{\beta}\right)$ and $\hat{\theta}_{1-\frac{\alpha-\delta}{2}}\left(x;\hat{\beta}\right)$, respectively. If our model for $p(\cdot)$ is correctly specified, Bonferroni's inequality implies that this interval covers $\theta$ with probability at least $1-\alpha$ in large samples.\footnote{Even in cases where we do not have an asymptotically normal estimator for $\beta,$ for example because we consider a fully nonparameteric model for $p(\cdot)$, given an initial level $1-\delta$ confidence set $CS_\beta$ for $\beta$ we can form a Bonferroni confidence set for $\theta$ as $\left[\inf_{\beta\in CS_\beta}\hat{\theta}_{\frac{\alpha-\delta}{2}}\left(X;\hat{\beta}\right),\sup_{\beta\in CS_\beta}\hat{\theta}_{1-\frac{\alpha-\delta}{2}}\left(X;\hat{\beta}\right)\right].$}

Applications

This section applies the results developed above to estimate the degree of selectivity in several empirical literatures.

\paragraph{Key identifying assumptions} The results of Section (ref) imply nonparametric identification of both $p(\cdot)$ and $\mu$. Using systematic replication studies, we identify $p(\cdot)$ based on asymmetries in the joint distribution of original and replication estimates. This approach is based on the assumption that selection for publication depends only on the original estimates and not on the replication estimates. This assumption is highly plausible by design in the two replication settings we consider.

Identification using meta-studies identifies $p(\cdot)$ by comparing the distribution of estimates with different associated standard errors across studies. This approach is based on the assumption that studies with different sample sizes on a given topic do not have systematically different estimands. While we cannot guarantee validity of this assumption by design, plausibility of this assumption is enhanced by our finding that it yields estimates very similar to the approach based on replication studies. It should also be noted that variants of this assumption are imposed in the vast majority of existing meta-studies.

\paragraph{Maximum likelihood estimation} The sample sizes in our applications are limited. For our main analysis, we specify parsimonious parametric models for both the conditional publication probability $p(\cdot)$ and the distribution $\mu$ of true effects across latent studies, which we then fit by maximum likelihood. Parametric specifications of the nonparametrically identified model lead to intuitive and tractable estimators. In the supplement we consider alternative, moment-based estimators which build on our identification arguments in Section (ref). These estimators are nonparametric in $\mu$ and yield similar results to the parametric specifications reported here.

We consider step function models for $p(\cdot),$ with jumps at conventional critical values, and possibly at zero. Since $p(\cdot)$ is only identified up to scale, we impose the normalization $p(z)=1$ for $z>1.96$ throughout. This is without loss of generality, since $p(\cdot)$ is allowed to be larger than $1$ for other cells.

We assume different parametric models for the distribution of latent effects $\Theta^*$, discussed case-by-case below. In our first two applications the sign of the original estimates is normalized to be positive.\footnote{The studies in these datasets consider different outcomes, so the relative signs of effects across studies are arbitrary. Setting the sign of the initial estimate in each study to be positive has the desirable effect of ensuring invariance to the sign normalization chosen by the authors of each study.} We denote these normalized estimates by $W=|Z|,$ and in these settings we impose that $p(\cdot)$ is symmetric.

\paragraph{Details and extensions} Details and further motivation for our specifications, as well as a specification for the model of Section (ref) which we use to develop specification tests, are discussed in Section (ref) of the supplement.

In the present section we assume that our identifying assumptions hold unconditionally for the samples considered. In Section (ref) of the Supplement, we explore robustness of our results to additional conditioning on covariates $W$, including year of first circulation and journal of publication. However, in no case do we reject our baseline (unconditional) specification at conventional significance levels. To check the robustness of our findings, we report additional empirical results based on further additional specifications in Section (ref) of the supplement. To check robustness to our parametric assumptions, we report estimates based on an alternative GMM estimation approach that does not rely on parametric specifications of $\mu$ in Section (ref) of the supplement. This alternative approach yields broadly similar estimates of $p(\cdot)$.

Economics laboratory experiments

Our first application uses data from a recent large-scale replication of experimental economics papers by camerer2016evaluating. The authors replicated all 18 between-subject laboratory experiment papers published in the American Economic Review and Quarterly Journal of Economics between 2011 and 2014.\footnote{In their supplementary materials, camerer2016evaluating state that “To be part of the study a published paper needed to report at least one significant between subject treatment effect that was referred to as statistically significant in the paper.” However, we have reviewed the issues of the American Economic Review and Quarterly Journal of Economics from the relevant period, and confirmed that no studies were excluded due to this restriction.} Further details on the selection and replication of results can be found in camerer2016evaluating, while details on our handling of the data are discussed in the supplement.

A strength of this dataset for our purposes, beyond the availability of replication estimates, is the fact that it replicates results from all papers in a particular subfield published in two leading economics journals over a fixed period of time. This mitigates concerns about the selection of which studies to replicate. Moreover, since the authors replicate 18 such studies, it seems reasonable to think that they would have published their results regardless of what they found, consistent with our assumption that selection operates only on the initial studies and not on the replications.

A caveat to the interpretation of our results is that camerer2016evaluating select the most important statistically significant finding from each paper, as emphasized by the original authors, for replication. This selection changes the interpretation of $p(\cdot)$, which has to be interpreted as the probability that a result was published and selected for replication. In this setting, our corrected estimates and confidence intervals provide guidance for interpreting the headline results of published studies.

\paragraph{Histogram} Before we discuss our formal estimation results, consider the distribution of originally published estimates $W=|Z|$, shown by the histogram in the left panel of Figure (ref). This histogram suggests of a large jump in the density $f_W(\cdot)$ at the cutoff $1.96$, and thus of a corresponding jump of the publication probability $p(\cdot)$ at the same cutoff; cf. the discussion in Section (ref). Such a jump is confirmed by both our replication and meta-study approaches.

\paragraph{Results from replication specifications}

figure[figure omitted — 628 chars of source]

The middle panel of Figure (ref) plots the joint distribution of $W,$ $W^r$ in the replication data of camerer2016evaluating, using the same conventions as in Figure (ref). To estimate the degree of selection in these data we consider the model

align*[align* omitted — 136 chars of source]

This assumes that the absolute value of the true effect $\Theta^*$ follows a gamma distribution with shape parameter $\kappa$ and scale parameter $\lambda.$ This nests a wide range of cases, including $\chi^2$ and exponential distributions, while keeping the number of parameters low. Our model for $p(\cdot)$ allows a discontinuity in the publication probability at $|Z|=1.96,$ the critical value for a 5% two-sided z-test. Fitting this model by maximum likelihood yields the estimates reported in the left panel of Table (ref). Recall that $\beta_p$ in this model can be interpreted as the publication probability for a result that is insignificant at the 5% level based on a two-sided z-test, relative to a result that is significant at the 5% level. These estimates therefore imply that significant results are more than thirty times more likely to be published than insignificant results. Moreover, we strongly reject the hypothesis of no selectivity, $H_0:\beta_p=1$.

To test the validity of our baseline assumption of selection on $z$ and not on $\theta$, $p(z,\theta)=p(z),$ we calculate a score test using a model discussed in Section (ref) of the supplement. This yields a p-value of 0.53, so we find no evidence that the assumption $P\left(D=1|Z^*,\Theta^*\right) = p(Z^*)$ imposed in our baseline model is violated.

table[table omitted — 1,134 chars of source]

\paragraph{Results from meta-study specifications}

While the camerer2016evaluating data include replication estimates, we can also apply our meta-study approach using just the initial estimates and standard errors. Since this approach relies on additional independence assumptions, comparing these results to those based on replication studies provides a useful check of the reliability of our meta-analysis estimates.

We begin by plotting the data used by our meta-analysis estimates in the right panel of Figure (ref). We consider the model

align*[align* omitted — 169 chars of source]

noting that $\Theta^*$ is now the mean of $X^*,$ rather than $Z^*,$ and thus that the interpretation of $(\tilde\kappa,\tilde\lambda)$ differs from that of $(\kappa,\lambda)$ in our replication specifications. Fitting this model by maximum likelihood yields the estimates reported in the right panel of Table (ref). Comparing these estimates to those in the left panel, we see that the estimates from the two approaches are similar, though the metastudy estimates suggest a somewhat smaller degree of selection. Hence, we find that in the camerer2016evaluating data we obtain similar results from our replication and meta-study specifications.

\paragraph{Bias correction}

figure[figure omitted — 965 chars of source]

To interpret our estimates, we calculate our median-unbiased estimator and confidence sets based on our replication estimate $\beta_p=.029.$ Figure (ref) plots the median unbiased estimator, as well as the original and adjusted confidence sets (with and without bonferroni corrections), for the 18 studies included in camerer2016evaluating. Considering the first panel, which plots the median unbiased estimator along with the original and replication estimates, we see that the adjusted estimates track the replication estimates fairly well but are smaller than the original estimates in many cases. The second panel plots the original estimate and conventional 95% confidence set in blue, and the adjusted estimate and 95% confidence set in black. As we see from this figure, even without Bonferroni corrections twelve of the adjusted confidence sets include zero, compared to just two of the original confidence sets. Hence, adjusting for the estimated degree of selection substantially changes the number of significant results in this setting.

Psychology laboratory experiments

Our second application is to data from open2015estimating, who conducted a large-scale replication of experiments in psychology. The authors considered studies published in three leading psychology journals, Psychological Science, Journal of Personality and Social Psychology, and Journal of Experimental Psychology: Learning, Memory, and Cognition, in 2008. They assigned papers to replication teams on a rolling basis, with the set of available papers determined by publication date. Ultimately, 158 articles were made available for replication, 111 were assigned, and 100 of those replications were completed in time for inclusion in open2015estimating. Replication teams were instructed to replicate the final result in each article as a default, though deviations from this default were made based on feasibility and the recommendation of the authors of the original study. Ultimately, 84 of the 100 completed replications consider the final result of the original paper.

As with the economics replications above, the systematic selection of results for replication in open2015estimating is an advantage from our perspective. A complication in this setting, however, is that not all of the test statistics used in the original and replication studies are well-approximated by z-statistics (for example, some of the studies use $\chi^2$ test statistics with two or more degrees of freedom). To address this, we limit attention to the subset of studies which use z-statistics or close analogs thereof, leaving us with a sample of 73 studies. Specifically, we limit attention to studies using z- and t-statistics, or $\chi^2$ and F-statistics with one degree of freedom (for the numerator, in the case of F-statistics), which can be viewed as the squares of z- and t-statistics, respectively. To explore sensitivity of our results to denominator degrees of freedom for t- and F-statistics, in the supplement we limit attention to the 52 observations with denominator degrees of freedom of at least 30 in the original study and find quite similar results.

\paragraph{Histogram} The distribution of originally published estimates $W$ is shown by the histogram in the left panel of Figure (ref). This histogram is suggestive of a large jump in the density $f_W(\cdot)$ at the cutoff $1.96$, as well as possibly a jump at the cutoff $1.64$, and thus of corresponding jumps of the publication probability $p(\cdot)$ at the same cutoffs. Such jumps are again confirmed by the estimates from both our replication and meta-study approaches.

\paragraph{Results from replication specifications}

figure[figure omitted — 630 chars of source]

The middle panel of Figure (ref) plots the joint distribution of $W,$ $W^r$ in the replication data of open2015estimating. We fit the model

align*[align* omitted — 174 chars of source]

This model again assumes that the absolute value of the true effect $|\Theta^*|$ follows a gamma distribution across latent studies. Given the larger sample size, we consider a slightly more flexible model than before and allow discontinuities in the publication probability at the critical values for both 5% and 10% two-sided z-tests.

table[table omitted — 1,075 chars of source]

Fitting this model by maximum likelihood yields the estimates reported in the left panel of Table (ref). These estimates imply that results that are significantly different from zero at the 5% level are over a hundred times more likely to be published than results that are insignificant at the 10% level, and nearly five times more likely to be published than results that are significant at the 10% level but insignificant at the 5% level. We strongly reject the hypothesis of no selectivity.

A score test of the null hypothesis $p(z,\theta)=p(z)$ yields a p-value of 0.42. Thus, we again find no evidence that the assumption $P\left(D=1|Z^*,\Theta^*\right) = p(Z^*)$ imposed in our baseline model is violated.

\paragraph{Results from meta-study specifications} As before, we re-estimate our model using our meta-study specifications, and plot the joint distribution of estimates and standard errors in the right panel of Figure (ref). Fitting the model yields the estimates reported in the right panel of Table (ref). As in the last section, we find that the meta-study and replication estimates are broadly similar, though the meta-study estimates again suggest a somewhat more limited degree of selection.

\paragraph{Bias corrections}

figure[figure omitted — 423 chars of source]

To interpret our results, we plot our median-unbiased estimates based on the open2015estimating data in Figure (ref). We see that our adjusted estimates track the replication estimates fairly well for studies with small original z-statistics, though the fit is worse for studies with larger original z-statistics. Our adjustments again dramatically change the number of significant results, with 62 of the 73 original 95% confidence sets excluding zero, and only 28 of the adjusted confidence sets (not displayed) doing the same.

\paragraph{Approved replications} Gilbertetal2016 argue that the protocols in some of the open2015estimating replications differed substantially from the initial studies. To explore robustness with respect to this critique, in the supplement we report results from further restricting the sample to the subset of replications which used protocols approved by the original authors prior to the replication. Doing so we find roughly similar estimates, though the estimated degree of selection is smaller.

Effect of minimum wage on employment

Our third application uses data from wolfson201515, who conduct a meta-analysis of studies on the elasticity of employment with respect to the minimum wage. In particular, wolfson201515 collect analyses of the effect of minimum wages on employment that use US data and were published or circulated as working papers after the year 2000. They collect estimates from all studies fitting their criteria that report both estimated elasticities of employment with respect to the minimum wage and standard errors, resulting in a sample of a thousand estimates drawn from 37 studies, and we use these estimates as the basis of our analysis. For further discussion of these data, see wolfson201515.

Since the wolfson201515 sample includes both published and unpublished papers, we evaluate our estimators based on both the full sample and the sub-sample of published estimates. We find qualitatively similar answers for the two samples, so we report results based on the full sample here and discuss results based on the subsample of published estimates in the supplement. We define $X$ so that $X>0$ indicates a negative effect of the minimum wage on employment.

\paragraph{Histogram} Consider first the distribution of the normalized estimates $Z$, shown by the histogram in the left panel of Figure (ref). This histogram is somewhat suggestive of jumps in the density $f_Z(\cdot)$ around the cutoffs $-1.96$, $0$, and $1.96$, and thus of corresponding jumps in the publication probability $p(\cdot)$ at the same cutoffs; these jumps seem less pronounced than in our previous applications, however.

\paragraph{Results from meta-study specifications}

figure[figure omitted — 458 chars of source]

For this application we do not have any replication estimates, and so move directly to our meta-study specifications. The right panel of Figure (ref) plots the joint distribution of $X$, the estimated elasticity of employment with respect to decreases in the minimum wage, and the standard error $\sigma$ in the wolfson201515 data.

As a first check, we run meta-regressions as discussed in section (ref), clustering standard errors by study. A regression of $X$ on $\sigma$ yields a slope of $0.408$ with a standard error of $0.372$. A regression of $Z$ on $1/\sigma$ yields an intercept of $0.343$ with a standard error of $0.283$. Both of these estimates suggest selection favoring results finding a negative effect of minimum wages on employment, but neither allows us to reject the null of no selection at conventional significance levels.

We next consider the model

align*[align* omitted — 244 chars of source]

Since the data are not sign-normalized, we model $\Theta^*$ using a t distribution with degrees of freedom $\tilde\nu$ and location and scale parameters $\bar\theta$ and $\tilde\tau,$ respectively. Unlike in our previous applications, we allow the probability of publication to depend on the sign of the z-statistic $X/\sigma$ rather than just on its absolute value. This is important, since it seems plausible that the publication prospects for a study could differ depending on whether it found a positive or negative effect of the minimum wage on employment. Recall that $X>0$ indicates a negative effect of the minimum wage on employment. Our estimates based on these data are reported in Table (ref), where we find that results which are insignificant at the 5% level are about 30% as likely to be published as are significant estimates finding a negative effect of the minimum wage on employment. Our point estimates also suggest that studies finding a positive and significant effect of the minimum wage on employment may be less likely to be published, but this estimate is quite noisy and we cannot reject the hypothesis that selection depends only on signficance and not on sign.

table[table omitted — 563 chars of source]

These results are consistent with the meta-analysis estimates of wolfson201515, who found evidence of some publication bias towards a negative employment effect, as well as the results of CardKrueger1995, who focused on an earlier, non-overlapping set of studies.

Since the studies in this application estimate related parameters, it is also interesting to consider the estimate $\bar\theta$ for the mean effect in the population of latent estimates. The point estimate suggests that the average latent study finds a small but statistically significant negative effect of the minimum wage on employment. This effect is about half as large as the “naive” average effect $\bar\theta$ we would estimate by ignoring selectivity, $.041$ with a standard error of $0.011$.

\paragraph{Multiple estimates} A complication arises in this application, relative to those considered so far, due to the presence of multiple estimates per study. Since it is difficult to argue that a given estimate in each of these studies constitutes the “main” estimate, restricting attention to a single estimate per study would be arbitrary. This somewhat complicates inference and identification.

For inference, it is implausible that estimate standard-error pairs $X_j, \sigma_j$ are independent within study. To address this, we cluster our standard errors by study.

For identification, the problem is somewhat more subtle. Our model assumes that the latent parameters $\Theta_i^*$ and $\sigma_i^*$ are statistically independent across estimates $i$, and that $D_i$ is independent of $(\Theta_i^*, \sigma_i^*)$ conditional on $X_i^*/\sigma_i^*$. It is straightforward to relax the assumption of independence across $i,$ provided the marginal distribution of $(\Theta_i^*, \sigma_i^*,X_i^*,D_i)$ is such that $D_i$ remains independent of $(\Theta_i^*, \sigma_i^*)$ conditional on $X_i^*/\sigma_i^*$. This conditional independence assumption is justified if we believe that both researchers and referees consider the merits of each estimate on a case-by-case basis, and so decide whether or not to publish each estimate separately. Alternatively, it can also be justified if the estimands $\Theta_i^*$ within each study are statistically independent (relative to the population of estimands in the literature under consideration). As discussed in Section (ref), however, if these assumptions fail our model is misspecified.

Deworming meta-study

Our final application uses data from the recent meta-study deworming2016 on the effect of mass drug administration for deworming on child body weight. They collect results from randomized controlled trials which report child body weight as an outcome, and focus on intent-to-treat estimates from the longest follow-up reported in each study. They include all studies identified by the previous review of Cochrane2015, as well as additional trials identified by Campbell2017. They then extract estimates as described in deworming2016 and obtain a final sample of 22 estimates drawn from 20 studies, which we take as the basis for our analysis. For further discussion of sample construction, see Cochrane2015, deworming2016, and Campbell2017. To account for the presence of multiple estimates in some studies, we again cluster by study.

\paragraph{Histogram} Consider first the distribution of the normalized estimates $Z$, shown by the histogram in the left panel of Figure (ref). Given the small sample size of 22 estimates, this histogram should not be interpreted too strongly. That said, the density of $Z$ appears to jump up at $0$, which suggests selection toward positive estimates.

\paragraph{Results from meta-study specifications}

figure[figure omitted — 451 chars of source]

The right panel of Figure (ref) plots the joint distribution of $X$, the estimated intent to treat effect of mass deworming on child weight, along with the standard error $\sigma$ in the deworming2016 data.

As a first check, we again run meta-regressions as discussed in Section (ref), clustering standard errors by study. A regression of $X$ on $\sigma$ yields a slope of $-0.296$ with a standard error of $0.917$. A regression of $Z$ on $1/\sigma$ yields an intercept of $0.481$ with a standard error of $0.889$. Neither of these estimates allows rejection of the null of no selection at conventional significance levels.

We next consider the model

align*[align* omitted — 166 chars of source]

where we constrain the the distribution of $\Theta^*$ to be normal and the function $p(\cdot)$ to be symmetric to limit the number of free parameters, which is important since we have only 22 observations. Fitting this model yields the estimates reported in Table (ref). The point estimates here suggest that statistically significant results are less likely to be included in the meta-study of deworming2016 than are insignificant results. \\

table[table omitted — 423 chars of source]

However, the standard errors are quite large, and the difference in publication (inclusion) probabilities between significant and insignificant results is itself not significant at conventional levels, so there is no basis for drawing a firm conclusion here. Likewise, the estimated $\bar\theta$ suggests a positive average effect in the population, but is not significantly different from zero at conventional levels.

In the supplement we report results based on alternative specifications which allow the function $p(\cdot)$ to be asymmetric. These specifications suggest selection against negative estimates.

Our findings here are potentially relevant in the context of the controversial debate surrounding mass deworming; see for instance wormwars2015. The point estimates for our baseline specification suggest that insignificant results have a higher likelihood of being included in deworming2016 relative to significant ones. In light of the large standard errors and limited robustness to changing the specification of $p(\cdot)$, however, these findings should not be interpreted too strongly.

Conclusion

This paper contributes to the literature in three ways. First, we provide nonparametric identification results for selectivity (in particular, the conditional publication probability) as a function of the empirical findings of a study. Second, we provide methods to calculate bias-corrected estimators and confidence sets when the form of selectivity is known. Third, we apply the proposed methods to several literatures, documenting the varying scale and kind of selectivity.

\paragraph{Implications for empirical research} What can researchers and readers of empirical research take away from this paper? First, when conducting a meta-analysis of the findings of some literature, researchers may wish to apply our methods to assess the degree of selectivity in this literature, and to apply appropriate corrections to individual estimates, tests, and confidence sets. We provide code on our webpages which implements the proposed methods for a flexible family of selection models.

Second, when reading empirical research, readers may wish to adjust the published point estimates and confidence sets along the lines discussed in Section (ref). Suppose for instance that for a given field publication probabilities increase considerably when estimates exceed the 5% significance threshold, but publication does not otherwise depend on findings. In that case, if reported effects are close to zero, or very far from zero (z-statistic bigger than 4, say), then these estimates can be taken at face value. In intermediate ranges, in particular for z-statistics around 2, magnitudes should be adjusted downwards.

It should be emphasized that we do not advocate adjusting publication standards to reflect our corrected critical values. If these cutoffs were to be systematically used in the publication process, this would simply entail an “arms race” of selectivity, rendering the more stringent critical values invalid again.

\paragraph{Optimal publication rules} One might take the findings in this paper, and the debate surrounding publication bias more generally, to indicate that the publication process should be non-selective with respect to findings. This might for instance be achieved by instituting some form of result-blind review. The hope would be that non-selectivity of the publication process might restore the validity (unbiasedness, size control) of standard inferential methods.

Note, however, that optimal publication rules may depend on results. Consider for instance a setting where policy decisions are made based on published findings, policy makers have a limited capacity to read publications, and journal editors maximize the same social welfare function as policy makers. In a stylized model of such a setting, detailed in Section (ref) of the supplement, we show that expected social welfare is maximized by publishing the results which allow policy makers to update the most relative to their prior beliefs. The corresponding publication rule favors the publication of surprising findings, thus violating non-selectivity. A more general theory of optimal publication is of considerable interest for future research.