Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
98,291 characters · 17 sections · 73 citation commands
Testing for Underpowered Literatures
\noindentKeywords: Deconvolution, Meta-Analysis, Statistical Power
\noindentJEL Codes: C12, C14, C90
\pagenumbering{arabic} \onehalfspacing
A key tradeoff in experimental design is balancing the risk of drawing an erroneous conclusion with project cost. Nearly all experiments published in top economics journals use $t$-tests to interpret their results. Sampling variance means that the $t$-test can fail to reject the null hypothesis of zero treatment effect even when there is in fact an effect of meaningful magnitude. Every experiment therefore runs the risk of a false negative. Collecting more data raises statistical power and reduces the risk of a false negative, but by an amount that depends on the true effect which is ex ante unknown. A central question faced by funders and researchers is when to focus resources on larger samples and when to direct funds toward other drivers of power, e.g. measurement quality McKenzie.
This paper provides research funders and meta-analysts with a statistical procedure to estimate how much larger the expected fraction of statistically significant $t$-scores would have been had every experiment in some population been counterfactually run with $c^2$ times the sample size where $c > 1$. This power gain is a number between zero and one that I call $\Delta_c$. If this quantity is large within a given scientific literature or funding initiative, then the rejection decisions of $t$-tests reported by that population of experiments are sensitive on average to sample size choices. In this case, the returns to collecting more data points could be relatively high.
To estimate $\Delta_c$ with the proposed method, the meta-analyst needs a dataset of $n$ $t$-scores reported by a sample of experiments. The method works by first finding the distribution of true intervention treatment effects that best fits a smoothed version of the empirical distribution of $t$-scores. The fitted distribution of true effects is then integrated to calculate an estimate of $\Delta_c$. I show that this procedure is consistent, asymptotically normal, and converges in a power of $n$.
The main contribution of this paper is that it allows the set of experiments being studied to be highly heterogeneous. The outcomes, designs, interventions, scales, and settings can vary across studies in any way. This is accomplished by removing all assumptions about the distribution of true intervention treatment effects that related meta-analyses have so far relied upon. The methods of PowerOfBias, Brunner, and others are only consistent when the population distribution of true intervention treatment effects has a specific shape. Restrictions of this kind need not hold when the literature being studied is, for example, a set of experiments united by a publisher or funder rather than by a type of intervention. This paper makes no assumptions of any kind about the distribution of true effects and so any kind of effect heterogeneity is allowed.
A second contribution of this paper is to make its method robust to simple forms of publication bias. Recent empirical evidence shows that statistically insignificant $t$-scores are less likely to be published in academic journals Franco1502,Andrews, Brodeur,bb, Elliott. Selective reporting can distort the distribution of reported $t$-scores and confound estimation. This paper addresses selective reporting by parameterizing a simple model of publication bias, estimating the model, and then reweighting the observed $t$-scores to remove publication bias. Accommodating this first stage requires that I use a particular smoothing method which results in a different estimator than those used to solve mathematically related {\it deconvolution} problems like CarrascoPaper.
An empirical estimate of $\Delta_c$ is useful to meta-analysts and funders because they may otherwise lack a way to evaluate how efficiently sample sizes are being chosen in practice. Sample size decisions are typically made under high uncertainty. If funders knew ex ante how quickly the statistical power of a proposed experiment responded to its sample size, they could optimally balance power and cost. But power also depends on how effective the treatment actually is, a quantity which is ex ante unknown. Grant-giving organizations try to address this challenge by requiring experimenters to collect samples large enough to guarantee at least 80% power to detect a true effect larger than some chosen threshold jpal_power_calcs. One of the main aims of these {\it power calculations} is to direct resources to projects that will only fail to detect an effect when that effect is truly small.
Power calculations are not guaranteed to achieve this goal in practice because they involve a great deal of guesswork. Statistical power depends on many unknown parameters besides the true effectiveness of the intervention, e.g. how heterogeneous treatment effects are across individuals. Even power calculations based on data from previous experiments are likely to systematically under-target sample sizes VU2024105868. If the power calculations are noisy or biased, then the choice of sample size made on their basis may trade off power and cost sub-optimally.
Even ex post it is difficult to determine what the consequences of a larger sample size would have been for a single experiment. This problem arises because it is not in general possible to infer how powerful an individual experiment was. Plugging the estimated treatment effect into the power function is known to be highly misleading posthocbad. In fact, precisely estimating the full distribution of true power over a literature is infeasible without very strong assumptions Fan91. Experimenters and funders therefore lack a rigorous way to assess how efficiently sample sizes are being chosen in practice. This paper addresses this need by showing how to estimate {\it average} power over a heterogeneous collection of studies under counterfactual sample sizes.
This paper applies its method to an empirical question of broad interest: how sensitive are randomized controlled trials (RCTs) published in top economics journals to counterfactual sample size increases? I use the data from bb which contains $t$-tests of main hypotheses reported by articles published in 25 top economics journals during 2015-2018. This is a very diverse set of experiments and was chosen to highlight the econometric contribution of this paper: a method robust to arbitrary heterogeneity in interventions, scales, settings, and designs. I estimate that counterfactually doubling the sample sizes of every RCT in that set would only increase the expected number of $t$-scores clearing the critical value of 1.96 by 7.2 percentage points with a standard error of 2.5.
This power gain is small in comparison to several benchmarks. First, I estimate that doubling the sample size of all non-RCTs in the same dataset would increase average power by 17.3 percentage points which is significantly larger. As a second benchmark, consider a hypothetical literature where every experiment is adequately powered to detect the true effect. By conventional standards, a literature where every experiment were powered at exactly 80% power would meet this criterion jpal_power_calcs. For such a hypothetical literature, we can calculate that doubling every sample size would increase power by 17.8 percentage points, which is significantly larger than the estimate for economics RCTs.
This paper constructs a third benchmark using data from the Many Labs systematic replication project ManyLabs. In this project, 36 laboratories each independently attempted to replicate 11 published effects from laboratory psychology. We would expect these replication experiments to be very well powered because they were designed using previous experimental data and the consequences of failure to detect an already published effect could be great. Yet I cannot reject the null hypothesis that $\Delta_c$ is the same for both RCTs in economics and the Many Labs replications. This means that RCTs are on average about as sensitive to sample size as replication experiments that were designed using a great deal of prior knowledge. The non-rejection is not simply from high uncertainty because we can indeed reject the null that $\Delta_c$ is the same for non-RCTs in economics vs replications from Many Labs. I also leverage the fact that the Many Labs experiments are replications of one another to estimate power gain conditional on the in-sample true effects. The conditional Many Labs power gain is precisely estimated at $7.8$ percentage points with a standard error of $0.6$ percentage points and is also statistically indistinguishable from the $\Delta_c$ for economics RCTs.
The takeaway from the application is that randomized trials in economics are on average relatively insensitive to doubling their sample sizes.\footnote{This is not to say that individual reported results do not need to be confirmed in replication. Rather this application says that had the original sample sizes been larger, replicability would not have improved much.} Power calculations appear to be surprisingly effective in practice---despite requiring guesswork {ex ante} and being impossible to directly verify {ex post}. Funders looking to improve power could consider alternative reforms instead. McKenzie argues that many field experiments could raise power substantially without collecting more data points by improving measurement quality and compliance. This paper's evidence suggests that the gains from making economics RCTs bigger is on average very modest and therefore quality improvements may deserve more attention.
This paper is organized as follows. Section (ref) discusses the contributions of this paper to existing literatures. Section (ref) sets up the problem without publication bias and explains why the deconvolution approach is needed. Section (ref) shows identification and Section (ref) proposes an estimator and derives its rate of consistency. Section (ref) introduces publication bias. Section (ref) shows asymptotic normality and discusses inference. Section (ref) uses simulations to recommend tuning parameters that yield high coverage of the confidence intervals for a variety of DGPs. This section also shows that coverage rates usually remain high when the $t$-score is only approximately normal. Section (ref) presents two empirical applications and Section (ref) concludes. The Online Appendix provides a key robustness check.
This paper contributes to two literatures. The first is an applied literature that empirically studies average statistical power over a population of experiments under very strict assumptions. PowerOfBias estimates median power in empirical economics under the assumption that papers can be sorted into groups within which every study is estimating the same true effect. A similar assumption is used by DellaVigna, bundock, and Ferraro. Other methods instead assume that the distribution of true effects has a specific shape, e.g. a Gamma distribution or a finite mixture model with fixed means Brunner,Sotola,zcurve2,NBERw31666. Assumptions like these need not hold in practice and are virtually impossible to check.
My main methodological contribution to this literature is to remove all assumptions about the distribution of true effects. This relaxation is important in theory because it produces robust results, allows true nulls to have positive probability, and accommodates outlier treatment effects that are of interest to the theory of experimental design ABfattails. These assumptions matter in practice because my nonparametric method yields substantively different results than the mixture model of zcurve2 in this paper's empirical application.
This paper also contributes on a conceptual level by suggesting that meta-analysts think of power “on the margin" and not just “in levels." This means that we should think of a collection of studies as “too small" when average power against true effects is easy to increase by growing the experiments, i.e. $\Delta_c$ is large. While many meta-studies have identified literatures where the level of power is low, this paper suggests refining this search and looking for places where power is easy to raise neuroPowerIoannidis,PowerOfBias. This has direct implications for research funding and design. When the returns to increasing sample size are low, other improvements in measurement and outcome choice might offer more fruitful solutions to power problems McKenzie.
This paper studies populations of experiments with arbitrary heterogeneity and therefore avoids specifying an experimenter utility function. It is not clear how to choose a function that maps effect magnitudes into welfare when studies can vary in their treatments, outcomes, scales, and settings---especially when not all studies are program evaluations. For this reason many analyses use the error rate of the hypothesis rejection decision as their welfare concept, e.g. neuroPowerIoannidis,Franco1502,Head,PowerOfBias,doi:10.1126/science.aac4716,young_ri,Brodeur,bb,Elliott. This paper estimates $\Delta_c$ because is interpretable even when the meta-sample encompasses a diverse set of interventions and can be used to direct research funding towards where it can improve power most. This estimand is relevant to, e.g. a grant-giver evaluating whether a particular funding initiative would have found more results if it had been given more resources.
The second related literature is technical. The estimation problem studied in this paper belongs to a class of problems called {\it deconvolutions} which aim to “de-blur" a smooth probability density. Deconvolutions are a classic problem Fan91,CARRASCO2Handbook,Meister,CarrascoPaper,handbookinverseproblems,Koenker03042014. But in the presence of publication bias standard methods will be inconsistent. There is growing evidence that statistically insignificant $t$-scores are less likely to be reported in economics publications Chrestensen, Andrews,Havranek24. Such omissions can create a discontinuity in the density of $t$-scores which is problematic for deconvolution caliperOriginal,Kudrinjmp.
Addressing publication bias is not as simple as concatenating deconvolution with publication bias removal because the two steps interact. This paper proposes a new deconvolution step in order to manage such interactions. The method begins with the same singular value decomposition as CarrascoPaper but uses a different method of regularization---spectral cutoff---that prevents uncertainty about the extent of selective reporting from magnifying the regularization bias. This choice of smoothing method yields a new estimator that converges faster and interacts minimally with publication bias removal.
Some recent meta-analyses identify publication bias via the joint distribution of standard errors and point estimates Andrews,Duval2000,Havranek24,vu2024pb. This strategy is not appropriate for my setting because it requires that the true standard errors and true effects are not correlated. In large meta-samples that encompass many interventions this independence breaks down because each researcher's choice of experimental design may be driven by knowledge about their true effect. This paper identifies selective reporting via the $t$-ratio like bb and Elliott.
First define a key piece of notation. Let $\varphi(z)$ denote the probability density function of the standard normal distribution. Adding a subscript $\varphi_{\sigma^2}(z)$ denotes the density of the normal distribution with variance $\sigma^2$. No subscript means that the variance is unity.
Consider a population of experiments. Each experiment studies a unique intervention with its own treatment effect $b \in \mathbb{R}$. The treatment effect $b$ is unobserved, but the experimenter estimates it with an estimator $\hat{b}$ that is unbiased and normally distributed: $\hat{b}\:|\:b\sim N\left(b,\sigma^2\right)$. The experimenter knows the standard error $\sigma$ and summarizes the evidence against the null hypothesis of zero treatment effect by reporting the $t$-score: $T = \frac{\hat{b}}{\sigma}$.\footnote{In practice, $T|h$ is only approximately normal. In theoretical meta-analysis it is common to ignore these concerns use the normal distribution regardless Elliott. In this paper, realistic deviations from normality are unlikely to matter much. Section (ref) presents a set of simulations where the numerator of the $t$-score is a sample mean of log-normals and another set where $t$-score is $t(30)$. In both cases the decline in coverage rates is small. } Let the random variable $h$ be called the “true effect" and define it as: $h \equiv \frac{b}{\sigma}$. The distribution of $T$ can now be described using $h$ only (making further references to $b,\sigma$ unnecessary). $T$ is conditionally normally distributed centered on $h$ with unit variance: $ T\: |\: h \sim N(h,1) $. Its conditional probability density function $ f_{T|h}\left(t\:|\:h\right)$ is the following:
$T$ can be interpreted as the test statistic of a two-sided $t$-test of the null hypothesis that $h= 0$. The power of a size-$\alpha$ test run by an individual experiment is the probability that $|T|$ exceeds the critical value $CV(\alpha)$ conditional on the true effect $h$. I suppress the dependence of the critical value on $\alpha$ for ease of notation. I call this conditional probability the {\it conditional power}. Conditional power can be written in terms of an integral over the normal density.
Since $h$ is unobserved, the meta-analyst cannot condition on it. So conditional power is always unknown. In practice this means that it is not possible to recover the power of any {\it individual} experiment by plugging its reported $t$-score into the power function posthocbad. To see why, notice that Jensen's Inequality guarantees that even though $E[T\:|\: h]=h$, nevertheless $\mathbb{E}\left[\varphi(t-T) \:|\: h\right] \neq \varphi(t-h)$.\footnote{Plugging the t-score into the power function also cannot be justified by invoking consistency of the point estimate because the t-score never converges in probability to a number.} Fortunately, the meta-analyst does not need to know the true power of any individual experiment. Instead the meta-analyst wishes to know the expected statistical power of an experiment {\it randomly drawn} from the population. I call this expectation the “unconditional power."
Unconditional power is defined as the expectation of conditional power over $h$. Unconditional power therefore depends crucially on the true probability distribution $\Pi_0$ of $h$. This paper will not require any restrictions on $\Pi_0$ of any kind---not even regularity conditions. This means that $\Pi_0$ could in principle be any mixture of discrete and continuous distributions and the moments of $h$ need not necessarily exist. This level of generality is necessary because we must accommodate the possibility of true nulls, i.e. probability mass at $h=0$.
Fubini's Theorem allows unconditional power to be expressed as an integral in Equation (ref).
Experimenters can influence the statistical power of their experiments by choosing the sample size. Increasing the sample size of an experiment will increase its conditional power whenever $h\neq 0$. The meta-analyst wishes to learn how much larger unconditional power would have been had the sample sizes of every experiment in the population been counterfactually increased while holding the distribution of true effects constant.
Define the random variable $T_c$ as a $t$-score randomly drawn from a counterfactual population where every experiment has been run at $c^2$ times the actual sample size while holding $\Pi_0$ constant.\footnote{In other words, this paper considers situations where there is an opportunity to collect more data without changing the treatment effect--i.e. the experimenter can sample more individuals without altering the sampling frame. This is often possible in practice given appropriate funding.} Since it is counterfactual, no draw of $T_c$ is ever actually observed. Multiplying the sample size by $c^2$ shrinks the standard error and grows $h$ by a factor of $c$. Conditional on the true effect, $T_c$ is normally distributed but with a larger mean, i.e. $T_c\:|\:h\: \sim N(ch,1)$. The unconditional counterfactual power is defined as the probability that $T_c$ exceeds the critical value. This can be expressed as the following double integral.
This paper proposes a method to determine whether the unconditional power of a given population of experiments is sensitive to sample size increases. High sensitivity represents an “opportunity missed" because there are in expectation many false negatives that could have been avoided had the experiments been larger. To make this idea precise, define the estimand $\Delta_{c}$ as the power gain resulting from increasing every sample size in the population of experiments by a factor of $c^2$.
$\Delta_c$ will be the estimand throughout this paper. Example (ref) and Figure (ref) illustrate why I interpret a large value of $\Delta_c$ as an indicator of an underpowered literature.
To identify and estimate $\Delta_c$, we will need to draw on the classic theory of deconvolution. To motivate the use of this mathematical framework, I pause here to briefly explain why counterfactual power cannot just be estimated using the simple and intuitive approaches that often come to mind. The fundamental difficulty is that power depends on how large the true effects tend to be, which is challenging to learn.
$T$ can be interpreted as the sum of $h$ plus an independent “noise" random variable $Z$ that has the normal distribution: $T=h+Z$. While it is not possible to estimate the power of an individual study conditional on $h$, it is straightforward to estimate the unconditional power at the status quo sample sizes in a simple way: just compute the fraction of $t$-scores that exceed the critical value:
Unfortunately, this method cannot be extended to estimate {\it counterfactual} power (or $\Delta_c$). That is, multiplying each $t$-score by $c$ and computing the fraction of those over $1.96$ will not be informative about what power would have been under larger sample sizes. The equation below shows that the probability limit of this procedure does not equal counterfactual power.
To see how consequential this gap is, consider an example literature where the true effect $h$ is equal zero for all studies. Here, power will be equal to size no matter how much the sample size grows. Nevertheless, the simple estimator in Equation (ref) will say that doubling sample sizes raises power to 16.7%! Why does this simple procedure fail? It is based on the idea that increasing the sample size by $c^2$ would decrease the denominator of the $t$-score (the standard error) by a factor of $c$. While this is true, collecting more data will also make the point estimate less volatile, pulling the numerator of the $t$-score towards the true effect of zero. Thus, while increasing the sample size of the experiment will decrease the denominator of $T$, it also changes the numerator as well.
The problem of learning $\Delta_c$ cannot be solved so easily because $\Delta_c$ depends on the unknown distribution $\Pi_0$ of $h$. In order to estimate $\Delta_c$ without any assumptions about $\Pi_0$, we need the convolutional framework developed in the following pages.
We saw above that $T$ is the sum of $h$ plus independent normal “noise" $Z$. The sum of two independent random variables is called a {\it convolution} and an operator that maps the distribution of $h$ into the distribution of $h$ plus noise is called a {\it convolution operator}. It will be useful to express the distributions of $T$ and $T_c$ in terms of convolution operators. The first step is to write down the densities of $T$ and $T_c$ as expectations over the distribution of true effects $\Pi_0$.
Both densities are the outcomes of closely related mappings of $\Pi_0$. Consider the operator $K_{\sigma^2}$ below that maps the distribution of $h$ to the density of $h+\sigma Z$ where $Z$ is a standard normal random variable independent of $h$ where $\sigma > 0$. This maps any probability distribution to a probability density.
A helpful fact about normal convolutions is that they can be decomposed. Lemma (ref) says that adding normal noise of unit variance is equivalent to adding normal noise of variance $c^{-2}$ and then adding further independent normal noise of variance $1-c^{-2}$.
The decomposition in Lemma (ref) is not new but since it is vital to all analysis going forward, its proof is verified in Appendix \textcolor{blue}{ (ref)}. It is immediate to see that $f_T = K_{1-c^{-2}}K_{c^{-2}}\Pi_0$. Lemma (ref) expresses the estimand $\Delta_c$ in terms of $K_{c^{-2}}\Pi_0$ as well. The proof is in Appendix (ref). The motivation for this decomposition is that $K_{c^{-2}}\Pi_0$ turns out to be much easier to estimate than $\Pi_0$ itself because it has already been smoothed out. The next section shows why this estimand is the case.
The meta-analyst observes $n$ draws of $T$ and wishes to estimate $\Delta_c$. This section will start by showing identification when $\Pi_0$ is continuous and has a probability density function $\pi_0$ with bounded height and there is no publication bias. Then we will see that the identification result in fact holds for all probability distributions $\Pi_0$---even those that are not continuous. Publication bias will be introduced in Section (ref) .
For now, assume that $h$ is continuous with PDF $\pi_0$ and that $\pi_0$ has finite height. The problem of recovering the PDF $\pi_0$ from $f_T$ is known to be severely ill-posed because very large changes in $\pi_0$ can result in very small changes in $f_T$. The intuition is that $f_T$ is a smoothed-out version of $\pi_0$. Since $\pi_0$ is not necessarily smooth itself, many of its “high-frequency" features are destroyed by convolution and are therefore difficult to recover from $f_T$. However, the problem of recovering $K_{c^{-2}}\Pi_0$ from $f_T$ is much better-posed because $K_{c^{-2}}\Pi_0$ has already lost its rapid oscillations.
This intuition can be formalized using the {\it singular value decomposition} (SVD). Intuitively, the singular value decomposition expresses the densities $K_{c^{-2}}\Pi_0$ and $\pi_0$ in terms of orthonormal basis polynomials with a one-to-one correspondence. The decompositions presented in this section are modified versions of the decompositions in CarrascoPaper and Wand1995KernelS. These decompositions were chosen because they will (later on) be made to accommodate publication bias and point masses in $\Pi_0$.
The first task is to precisely define the domain and range of the two convolution operators $K_{c^{-2}}$ and $K_{1-c^{-2}}$. The meta-analyst must start by choosing the scalar $\sigma_Y^2>0$. In the simulations and applications of this paper I always choose $\sigma_Y^2=1$ and never deviate from this choice. Define the following three Hilbert spaces of functions: $\mathcal{L}_{Y},\mathcal{L}_{X},\mathcal{L}_{W}$.
These three spaces are all very large. Each contains every bounded probability density function of a real-valued random variable. Equip each space with the following inner products:
These inner products induce norms: $\left|\left|\phi\right|\right|_Y^2 \equiv \langle \phi,\phi\rangle_Y $. The convolution operators can now be fully defined by specifying their domains: $ K_{c^{-2}} \::\: \mathcal{L}_W \to \mathcal{L}_X$ and $ K_{1-c^{-2}} \::\: \mathcal{L}_X\to \mathcal{L}_Y$. Both are compact linear operators and therefore must have { singular value decompositions}. In this case the singular value decomposition will express the outcome of a convolution as a weighted sum of known orthonormal polynomials where the weights are known to decay at an geometric rate. These decompositions are expressed below.
The singular values $\eta_j,\lambda_j$ are the sequences of scalars defined below. These decay geometrically fast to zero. The fast rate of decay implies that the problem of recovering $\pi_0$ from $K_1\Pi_0$ is ill-posed because components of $\pi_0$ with large $j$ play a small role in $K_1\Pi_0$.
The singular functions $\chi_j,\psi_j,\phi_i$ are the generalized Hermite Polynomials.\footnote{ The generalized Hermite Polynomials are $ He_j(t)=\frac{1}{\sqrt{j!}}\sum_{l=0}^{[j/2]}(-1)^l \frac{(2l)!}{2^ll!}\binom{j}{2l}t^{j-2l}$. The singular functions are scalings of the Hermite polynomials: $\chi_j(t) = He_j\left(\frac{t}{\sqrt{1+\sigma_Y^2}}\right)$, $\phi_j(t) = He_j\left(\frac{t}{\sqrt{1+\sigma_Y^2 -c^{-2}}}\right)$, and $\psi_j(t)=He_j\left(\frac{t}{\sigma_Y}\right)$.} I will use four properties of the Hermite polynomials. First, the polynomials are normalized so that, for example $\langle \chi_j,\chi_j\rangle_W = 1$. Second, each set forms a complete basis for its corresponding Hilbert space HermiteComplete. This means that for any two probability densities $\pi_1,\pi_2$, if $\langle \pi_1-\pi_2, \chi_j\rangle_{X} = 0$ for all $j$, then $\pi_1=\pi_2$ almost everywhere. Third, the polynomials are orthogonal, so $\langle \chi_j,\chi_k\rangle_W=0$ when $j\neq k$. Fourth, while these polynomials are themselves unbounded, they are uniformly bounded over all $t,j$ when multiplied by the kernels from their respective inner products HermiteBound. This means that:
Equation (ref) below expresses the density of $T$ as a linear combination of the Hermite polynomials. Notice that the singular values $\lambda_j$ and $\eta_j$ are forcing higher order polynomials to play a small role. The implication for the meta-analyst is that high-frequency information about $\pi_0$ is not easy to recover from $f_T$.
Lemma (ref) expresses the estimand $\Delta_c$ in terms of the singular values and Hermite polynomials.
The proof is in Appendix \textcolor{blue}{ (ref)}. Lemma (ref) illustrates why $\Delta_c$ is a fundamentally easier quantity to learn than $\pi_0$. The sequence of coefficients $\langle \chi_j,\pi_0 \rangle_W$ is sufficient for both the distribution of the data and for the estimand $\Delta_c$. Equation (ref) shows that the larger $j$ is, the more difficult it is to recover $\langle \chi_j,\pi_0 \rangle_W$ from $f_T$ since it is damped away by the rapidly decaying coefficients $\eta_j\lambda_j$. But, Lemma (ref) shows that when $\langle \chi_j,\pi_0 \rangle_W$ is given small weight in $f_T$, it is also given small weight in $\Delta_c$ because $\eta_j$ is decaying as well. So the pieces of $\pi_0$ that are hardest to recover thankfully also play the smallest role in $\Delta_c$.
The result in Lemma (ref) can be generalized to Theorem (ref) which places no restrictions on $\Pi_0$ at all. Theorem (ref) is an identification result because it expresses our estimand $\Delta_c$ in terms of the population distribution of the $t$-score under the factual sample sizes.
The proof is in Appendix (ref). The intuition for the argument is that for any $\Pi_0$ we can construct a sequence of densities $\pi_n$ (each with finite height) that converge weakly to $\Pi_0$. By the Portmanteau Theorem, the sequence of $\Delta_c$ for $\pi_n$ must converge to the $\Delta_c$ for $\Pi_0$. By dominated convergence, the infinite sum over the expectations over $\pi_n$ in Lemma (ref) converge to the infinite sum on the right hand side of Theorem (ref). Using Portmanteau a second time verifies that the limit of each sequence of expectations over $\pi_n$ equals the expectation over $\Pi_0$.
The meta analyst observes a sample of $n$ $t$-scores denoted $t_{i,k}$ indexed by $i\in \{1,\cdots n\}$ reported by $m$ studies indexed by $k \in \{1,\cdots m\}$. Let the number of t-scores per study be uniformly upper bounded by $B>0$. The $t$-scores are independent across studies, but can be dependent within a study. The $k$ subscript will sometimes be suppressed. For now there is no selective reporting.
Even though $\Pi_0$ is itself identified, estimating it is a severely ill-posed problem because the singular values $\lambda_j\eta_j$ decay rapidly to zero. This means that information about the high-frequency components of $\pi_0$ is severely attenuated in the distribution of $T$. The rate at which it is possible to estimate $\Pi_0$ depends on how we measure the difference between the estimated density and the true density. If we take this difference to be the $\mathcal{L}_\infty$ norm between CDFs, then the minimax rate is known to be logorithmic in $n$ Fan91. This is extremely slow. Fortunately, the meta-analyst only needs to estimate the features of $\Pi_0$ that matter for counterfactual power.
This paper proposes the estimator $\hat{\Delta}_{c,n}$ defined below which is the sample analogue of Theorem (ref).
This sample analogue replaces expectations with sample means. Instead of summing over all $j$, the estimator is regularized by summing only up to an integer $J_n$ that grows with $n$. This regularization method is called {\it spectral cutoff}.
The estimation error of $\widehat{\Delta}_{c,n}$ can be upper bounded by the two sums below. The first sum is random sampling error. The second sum is the deterministic “regularization bias" that is the consequence of halting the sum at $J_n$. Here $\lesssim$ means that the left side is upper bounded by a universal constant times the right side.
To see why regularization is necessary, consider the “sampling error" term. Since $\lambda_j$ is decaying to zero, if $J_n$ grows too quickly (or is infinite) then the variance of $ \hat{\Delta}_{c,n}$ can diverge. The expectation of the square of the sampling error can be bounded by studying the rate of decay of the $\lambda_j$ and by uniformly bounding the functions $\psi_j(t)\varphi_{\sigma_Y^2}(t)$ over all $j,t$ HermiteBound. This argument yields a bound on the sampling error.
The drawback of cutting the sum off at $J_n$ is regularization bias, i.e. approximation error. The bias can be bounded using the fact that $\chi_j(h)\varphi_{\sigma^2+1}(h)$ is uniformly bounded over $j,h$ HermiteBound. This means that the regularization bias is bounded by a geometric series which is the same order of its first summand.
The meta-analyst wishes to increase $J_n$ with the meta-sample size $n$ at a rate that makes $\widehat{\Delta}_{c,n}-\Delta_c$ converge in probability to zero as fast as possible. This means balancing regularization bias and sampling variance. Because we regularize with spectral cutoff, the rate-optimal choice of smoothing parameter does not depend on $c$. Theorem (ref) gives the optimized rate that equalizes the order of the bias and sampling error.
The proof is in Appendix \textcolor{blue}{ (ref)}. The rate of convergence is involved but it contains several insights. If the meta-analyst wishes to estimate $\Delta_c$ for a marginal increase in sample size, then they set $c=1+\epsilon$ with $\epsilon$ small. This makes $q\approx 1$ and the estimator approximately achieves the parametric rate $n^{-1/2}$. As the meta-analyst increases $c$, the rate of convergence slows down. This illustrates that $\Delta_c$ is harder to estimate when $c$ is large. The intuition is the following. The larger $c$ is, the more it matters whether $h$ is small and positive or exactly zero. Therefore for large $c$, the high-frequency components of $\Pi_0$ matter more for $\Delta_c$. Since the non-smooth components are harder to estimate, the rate of convergence slows down.
It is not realistic in practice to assume that every $t$-score computed by an experimenter is reported with equal probability. After $t$-scores are computed, only a subset may be published. There is growing evidence that $t$-scores in the social sciences are selected for publication partially on the basis of whether they cross certain significance thresholds Franco1502,Brodeur,bb,Andrews,Elliott. In this section I add publication bias to the problem. I show that $\Delta_c$ is still identified for a broad class of models of publication bias. Then I show how to consistently estimate $\Delta_c$ under the simplest canonical model.
This paper approaches publication bias by specifying a parametric model of selective reporting similar to Hedges1992. Here I provide a relatively weak sufficient condition for the model to be identified. Let $R$ be the event that the $t$-score $T$ was reported. Let the conditional probability of reporting be equal to:
where the form of the function $w_{\theta}(t)$ is known to the meta-analyst and the parameter $\theta_0$ is an unknown member of the compact set $\theta_0 \in \Theta\subseteq \mathbb{R}^K$. I make several assumptions about the publication bias model. Assumption (ref) below requires that $w_{\theta}(t)$ be bounded away from zero and from infinity. Without an assumption like this, $\Pi_0$ is not necessarily identified.
Under publication bias the meta-analyst observes only $T$ conditional on $R=1$. This is problematic because the expectations $E[\psi_j(T)\varphi_{\sigma_Y}(T)]$ from Theorem (ref) are no longer directly available in terms of the distribution of observed $t$-scores. Instead the population distribution available to the meta-analyst is the conditional distribution $T\:|\:R$. If the meta-analyst knew ${\theta_0}$, then to recover an expectation over $T$, they could take an expectation of $T|R$ times an appropriate weight. Lemma (ref) computes this weighting using Bayes' Theorem.
The proof is in Appendix \textcolor{blue}{ (ref)}. Combining this result with Theorem (ref) immediately yields Corollary (ref) below. This shows that if $\theta_0$ is identified, then $\Delta_c$ must also be identified. The remaining task is to identify $\theta_0$ itself.
This paper will identify $\theta_0$ using only the distribution of published $t$-ratios. I cannot identify $\theta_0$ via the joint distribution of published point estimates and standard errors as Andrews or Duval2000 do because this would require assumptions far too strong for this setting. The {\it funnel plot strategy} requires true effects and true standard errors to be uncorrelated. If this holds, then observing correlation in published point estimates and their standard errors reveals publication bias. This symmetry assumption is plausible in settings where every study is investigating a similar effect.
This paper allows each experiment to investigate a completely different effect. This flexibility allows the meta-analyst to study a set of experiments united by their funders or publishers and not necessarily by their topics. Full study heterogeneity can introduce dependence between true effects and their standard errors because researchers who know {\it ex ante} that they are probably investigating a small effect may choose a larger sample size or pre-specify a less robust but more precise estimator. Dependence can arise for other reasons as well. For instance, if some studies measure effects in dollars and others in euros, heterogeneity in the units will induce correlation between standard errors and true effects. Therefore this paper identifies $\theta_0$ via the $t$-curve which is not confounded by such dependence.
For $\theta_0$ to be identified from the distribution of $T|R$, publication bias must always distort the distribution of $t$-scores in a way that could never happen “naturally." I now add an assumption on the shape of $w_\theta(t)$ to this effect. In the absence of publication bias, $T$ must always have an infinitely differentiable probability density function because it is the outcome of a convolution with the normal distribution. Most models of publication bias specify that reporting decisions are made on the basis of statistical significance, i.e. on whether $T$ crosses certain thresholds. If publication bias introduces kinks, jumps, or non-smoothness into the density $f_{T|R}(t|R)$ (or its derivatives), then it is possible to recover the original smooth density. Assumption (ref) stipulates that publication bias “reveals itself" through breaks at a countable set of points. If these breaks take the form of discontinuities at traditional critical values they are sometimes called “Caliper Gaps" and they have been well studied by others caliperOriginal,Elliott,Kudrinjmp. However Assumption (ref) is more general because it envisions publication bias that reveals itself through any non-smoothness in the $t$-curve.
Theorem (ref) states that Assumption (ref) is sufficient for identification.
The proof is in Appendix (ref). The intuition for Theorem (ref) is the following. Suppose that the meta analyst has a guess $\theta$ for $\theta_0$ and attempts to remove the publication bias by reweighting the density $f_{T|R}$ by $1/w_{\theta}$ via Lemma (ref). Only by reweighting with the true $\theta_0$ can the meta-analyst remove all of the jump discontinuities in all the derivatives. So each $f_{T|R}$ is mapped to a unique $\theta$ and $\theta_0$ is identified from $f_{T|R}$. Since Corollary (ref) expresses $\Delta_c$ as a function of $f_{T|R}$ and $\theta_0$, then $f_{T|R}$ must identify $\Delta_c$ as well.
To estimate $\Delta_c$ under publication bias the meta-analyst must specify a parametric function $w_\theta(t)\: :\: \mathbb{R}\to \mathbb{R}^+$ that satisfies Assumptions (ref) and (ref). Many models are possible. For ease of exposition in this section I specify the most straightforward canonical model of publication bias from Andrews. Suppose that the probability of publication depends only on whether $|T|$ falls above the critical value $1.96$. The conditional probability ratio is equal to the scalar $\theta_0=\frac{\text{Pr}\left(R\:|\:|T|< 1.96\right)}{\text{Pr}\left(R\:|\:|T|\geq 1.96\right)}\in \left(\frac{1}{\overline{M}},\overline{M}\right)$. This means that the weighting function is:
This function satisfies Assumption (ref). Here publication bias reveals itself via a discontinuity in the t-curve at the traditional critical value of 1.96. This kind of “Caliper" discontinuity is well-studied caliperOriginal,Kudrinjmp. Equation (ref) expresses $\theta_0$ in terms of the jump discontinuity at $t=1.96$ in the density of published $t$-scores $f_{T|R}$.
The meta-analyst can construct a plug-in estimator $\widehat{\theta}_n$ by estimating the caliper jump in the histogram of published $t$-scores. To do this the meta-analyst chooses a bin width $\epsilon_n>0$ and takes the ratio between the number of t-scores within $\epsilon_n$ to the right vs to the left of $1.96$.
The meta-analyst can plug $\widehat{\theta}_n$ and the sample means into Corollary (ref) to obtain an estimator $ \widehat{\Delta}_{c,n}^{pb}$ for $\Delta_c$ under simple publication bias.
Theorem (ref) shows the consistency of $\widehat{\Delta}_{c,n}^{pb} $. Adding publication bias has slowed the rate of convergence down to $n^{-q/3}$. This happens because a small change in the histogram near a point in $1.96$ can change the weight that every $t_i$ gets in each of the sample averages in the sample objective.
In this section I show the asymptotic normality of $\hat{\Delta}_{n,c}^{pb}$ and derive a consistent variance estimator. The technical methods are standard (if lengthy) applications of Taylor linearization, constructing a triangular array, and checking the Lyapunov Condition.
I show that $\hat{\Delta}_c$ is asymptotically normal in the usual way by expressing the estimation error as a weakly dependent sample average using Taylor approximations. Regularization bias will affect the centering of the estimator. The centering term is non-negligible because $\hat{\Delta}_{n,c}^{pb}$ contains smoothing bias from two different sources. First, since only the first $J_n$ terms of the singular value decomposition are used, the rest of the terms are set to zero and this incurs regularization bias. Second, since the meta-analyst only observes a sample of $t$-scores, they estimate the discontinuities in $f_{T|R}$ by looking in a window of width $\epsilon_n$ on either side of each possible point of discontinuity. Since the density can change over this interval, smoothing bias is incurred here as well.
After centering properly, it is possible to linearize the sampling error of $\hat{\Delta}_{c,n}^{pb}$. Lemma (ref) below uses Taylor's Theorem several times to rewrite the estimator as a triangular array of sample means plus a vanishing term. The function $Z_{J_n,\epsilon_n} (t)$ is deterministic. Since it is lengthy to write down, its expression is relegated to the proof of the Lemma.
The proof is in Appendix \textcolor{blue}{ (ref)}. Next I show that the triangular array of sample means $ \frac{1}{n}\sum_{i=1}^n Z_{J_n,\epsilon_n} (t_i)$ is asymptotically normal. In order for this to guarantee that the estimator $ \widehat{\Delta}_{c,n}^{pb}$ is itself also normal, the sample means must dominate the Taylor residuals. To guarantee this I add Assumption (ref) which says that the estimator does not converge too fast.
To show asymptotic normality of the triangular array of sample sums I invoke a Lindeberg-Feller type Central Limit Theorem. The key step is to check the Lyapunov Condition. While verifying this condition is a common tactic in ill-posed problems, my argument is quite different than CarrascoPaper. Using the bounds on the magnitude and variance of the $Z_{J_n,\epsilon_n} $ from Lemma (ref), we can verify that Assumption (ref) is sufficient to guarantee the following Lyapunov Condition that uses the fourth moments. Since each observation of $T|R$ is identically distributed and independent of all but at most $D$ other observations, the dependence is weak. Theorem (ref) shows that the assumptions so far are enough to satisfy all of the hypotheses of Theorem 2.1 of Neumann. The formal argument is in Appendix \textcolor{blue}{ (ref)}.
Next I derive an estimator consistent for the variance of $\frac{1}{n}\sum_{i=1}^n Z_{J_n,\epsilon_n} (t_i)$. Equation (ref) expresses the variance of the sum as the sum of the covariances. Since two $t$-scores drawn from different studies are independent, the covariances are zero across studies. Let $\Lambda$ be the $n\times n$ block-diagonal matrix where the element $\Lambda_{ij}$ indicates whether $t_i$ and $t_j$ were reported in the same study. This gives us the following expression for the variance:
Since the functions $Z_{J_n,\epsilon_n} $ depend on the true $\theta_0$ and $\Pi_0$, the meta-analyst does not know them. But there is a sample version $\hat{Z}_{J_n,\epsilon_n} $ that the meta-analyst does observe and can be used for variance estimation. Let $F,\hat{F}_n$ denote the CDF and empirical CDF of $|T|$. Now $\hat{Z}_{J_n,\epsilon_n} $ can be expressed in the following way:
Define $\overline{Z}$ to be the sample mean of the $\hat{Z}_{J_n,\epsilon_n}(t_i)$. Theorem (ref) guarantees that variance estimation is consistent.
The proof is in Appendix \textcolor{blue}{ (ref)}. Theorem (ref) says that the meta-analyst can construct valid confidence intervals covering the centering sequence $\tilde{\Delta}_{c,n}$. Lemma (ref) showed that the centering sequence converges to $\Delta_c$ at the same rate as the variance decays. In theory, this could affect the coverage of the confidence intervals for $\Delta_c$ itself. A natural solution is to “undersmooth" or to change $J_n,\epsilon_n$ more quickly than the optimal rate in order to force $\tilde{\Delta}_{c,n}$ to converge to $\Delta_c$ faster than the variance decays. The simulations and empirical applications in this paper choose not to undersmooth but achieve nearly perfect coverage for the 95% confidence intervals in simulation nevertheless.
This section presents simulations showing near-perfect coverage of $\Delta_c$ by the 95% confidence intervals under a variety of circumstances. This section also illustrates how to set tuning parameters $\epsilon_n,$ $J_n$, and $\sigma_Y$ to yield this good coverage. To guarantee the optimized rates of convergence in Theorems (ref) and (ref) the meta-analyst sets $\epsilon_n = Cn^{-1/3}$ and $J_n = \log\left(Dn^{-1/3}\right)/\log\left(\sigma_Y^2/\left(1+\sigma_Y^2\right)\right)$ where $C,D>0$. This means that the meta-analyst's choice is actually over the triple of constants $\left\{C,D,\sigma_Y\right\}$. The theorems in the preceding sections provide no specific guidance on how these constants are to be chosen and we must turn to simulations. I recommend that the meta-analyst should always at least disclose results using the following tuning parameters: $C=2$, $D=10^{-4}$, and $\sigma_Y=1$. These choices of tuning parameters yield good confidence interval coverage of $\Delta_c$ in simulation for a very wide variety of data generating processes.
I generate simulated data in the following way. I use the simple model of publication bias from Section (ref) where $t$-scores are reported with probability $\theta_0$ if they do not clear $1.96$ and are reported with certainty otherwise. This matches the illustrative example from Andrews. In this section I set $\theta_0=0.9$ but the results are not sensitive to this. Table (ref) reports simulations for several choices of $\Pi_0$ where we expect coverage to be poor for one reason or another. Theory predicts that coverage could be low when the distribution $\Pi_0$ is not very smooth or if the density $f_T$ is close to zero or has steep slope at the critical threshold $1.96$. I report simulations against several different distributions $\Pi_0$, some of which are non-smooth.
The simulation results are reported in Table (ref). Here I compare performance for these five $\Pi_0$ under two different meta-sample sizes: a modest meta-sample of $50$ $t$-scores versus a large meta-sample of $500$ $t$-scores. The main message of Table (ref) is that coverage of $\Delta_c$ by the 95% confidence intervals is close to the nominal level under a broad variety of potentially problematic $\Pi_0$ for both large and small $n$.
In practice, $T|h$ is only approximately normally distributed. There are two reasons for this. First, the researcher must estimate the denominator of the $t$-score, which results in $T$ having a student-$t$ distribution. Table (ref) below shows that this is unlikely to be a concern when estimating $\Delta_c$. Here we repeat the Monte Carlo analysis in Table (ref), but simulate data with $T|h = h+t(30)$. Despite the small number of degrees of freedom, the coverage rates remain largely the same as before.
A more serious deviation from normality comes from the numerator of the $t$-score. When the effective sample size of the experiment is small, the Central Limit Theorem might not have fully “kicked in" yet. The Edgeworth Series tell us that this is primarily a concern when the outcome variable studied by the experiment is skewed or has excess kurtosis. To address this concern, we run the simulation in Table (ref). Here $T-h$ is the scaled and centered mean of $185$ i.i.d. log-normally distributed random variables. The number $185$ was chosen to match the median number of treatment clusters in the RCTs from the bb data.\footnote{ This is the median of a random sample of 100 of the RCTs in the dataset. A research assistant checked each paper individually to find the number of clusters. There was occasional ambiguity about the number of treatment clusters. In these cases we always used the smaller number. The median number of observations per study reported directly by bb is 5202.} The log-normal was chosen because it has substantial skew and excess kurtosis---making it a non-favorable distribution for the CLT. Nevertheless, coverage remains relatively high. These simulations suggest that the empirical estimates in the next section are unlikely to be significantly distorted by the inexactness of the normal approximation.
This section applies the methods proposed above to address an empirical question with important policy implications: Are randomized controlled trials (RCTs) published in top economics journals too small? RCTs are lauded as the “gold standard" of empirical evidence in social science because their design allows them to credibly control the rate of type-I errors. But what about type-II errors? If the conclusions reported by influential RCTs are sensitive to reasonable increases in sample size, then funders and researchers should run experiments with fewer treatment arms and allocate more resources toward data collection.
This is an empirical question and the answer will depend on the population of experiments under study. Here I study the experiments that have the most influence in academic economics: those published in top journals. The data source is bb. This meta-study examined the universe of 684 articles published in top economics journals during 2015-2018 of which 145 were RCTs. The data contain 21,740 test statistics of which 20,419 were $t$-scores that I can use. All test statistics corresponded to main hypotheses of interest and excluded covariates, placebo tests, etc. Every $t$-score was produced using one of four empirical methods: Randomized Controlled Trials (RCTs), Difference in Differences (DID), Discontinuity Designs (DD), and Instrumental Variables (IV). The $t$-scores are derounded.
I estimate $\Delta_c$ among RCTs and non-RCTs in Table (ref). The estimate in column 1 says that counterfactually doubling the sample sizes of every RCT published in top economics journals would only increase the expected number of $t$-scores clearing the critical value of 1.96 by 7.2 percentage points with standard error $2.5$.
I argue that 7.2 percentage points should be viewed as a small power gain and therefore that RCTs published in top economics journals are not very sensitive to sample size increases on average. This interpretation suggests that funders should sponsor more RCTs rather than fewer, larger ones. I defend the notion that 7.2 is a small gain by comparing it to three benchmarks.
As an initial benchmark, consider a literature where every experiment is run at exactly 80% power. This is the traditional threshold for an RCT to be considered “sufficiently powered" jpal_power_calcs. We can calculate that doubling every sample size would increase power by 17.8 percentage points. The confidence interval in Column 1 of Table (ref) does not come close to this value. So RCTs in practice are less than half as sensitive to sample size increases as a set of hypothetical “well-powered" experiments, and therefore must themselves be (I argue) very well powered.
As a second benchmark, I compare RCTs to natural experiments in the bb dataset. The data contains of 14171 t-scores reported in 559 articles that used observational methods to identify causal effects.\footnote{Twenty articles contained $t$-tests from both an RCT and an observational method.} Column 2 of Table (ref) reports that among tests conducted by non-RCTs, doubling sample sizes would increase power by $17.3$ percentage points, which is significantly larger than for RCTs $(p=0.00)$. Figure (ref) visualizes this contrast for a variety of sample size increases. It shows how power climbs faster for natural experiments than RCTs in the bb sample across the board.
The difference in sensitivity between RCTs and natural experiments is likely because experimenters who run RCTs choose their sample sizes. In contrast, observational researchers typically cannot choose the sample size of a natural experiment. Even if power calculations involve guesswork, they seem to still provide some useful information. Since randomized trials are already optimized in a way that quasi-experiments are not, it is not surprising that they have less to gain from sample size changes. The next subsection constructs a third benchmark that is more surprising.
This subsection presents a third benchmark to compare against the power gain of 7.2 percentage points estimated the previous subsection. I construct the benchmark using a second set of experiments that has two special properties: (i) we have good reason expect all of the experiments in this set to be very well powered and (ii) it is possible to construct a credible alternative estimate of the power gain as a robustness check.
The Many Labs systematic replication project ran a set of experiments with both of these special properties. ManyLabs recruited 36 independent research teams who each attempted to replicate 13 effects from laboratory psychology. Each team collected their own independent data and ran some or all of the thirteen experiments on their respective samples. Every sample contained at least 79 participants and many contained far more for a total of 6344. Eleven of the experiments were analyzed using $t$-tests of the equality in means between a treatment group and a control group and I will limit my analysis to these. This yields a meta-sample of 385 $t$-scores. The aim of all of the experiments run by Many Labs was to replicate effects that had been published in top psychology journals and to investigate how consistently effects could be replicated across study sites.
The key special property of replication projects is that the sample size of every experiment was chosen on the basis of data from a previous published experiment that studied the same effect. This means that we can expect the Many Labs experiments to be on average very well powered and to have a small $\Delta_c$ relative to what can reasonably be achieved in practice. In Table (ref) the 95% confidence interval contains the power gain of 7.2 percentage points from economics RCTs. This means that I cannot reject the null hypothesis that $\Delta_c$ for Many Labs is different than $\Delta_c$ for RCTs published in top economics journals $(p=0.17)$. That is to say, economics RCTs are not significantly poorer powered than laboratory replications that we ex ante expect to be very well powered. This is not simply a consequence of wide confidence intervals or high uncertainty because $\Delta_c$ for Many Labs is significantly different from $\Delta_c$ for non-RCTs in the bb data.
To show that the results are not sensitive to reasonable choices of the tuning parameters, Table (ref) presents four specifications. In columns 1 and 2 the tuning parameters $J_n,\epsilon_n$ are scaled by the number of experimental study sites. In columns 3 and 4 the tuning parameters are instead scaled by the total number of $t$-scores. In columns 1 and 3 I use the recommended values of the constants $C,D$ from Table (ref). In columns 2 and 4 I vary these choices to show that the confidence sets do not change much even when the changes to $C,D$ are large.
The Online Appendix presents another key robustness check. There I exploit a second key feature of the Many Labs setting: each laboratory was plausibly testing the same set of hypotheses. This special circumstance makes it possible to construct an alternative estimator that is much more precise than $\hat{\Delta}_c$ in the same spirit as PowerOfBias and bundock without imposing any additional assumptions. I estimate that doubling the sample size of every Many Labs replication (this time conditional on the hypotheses themselves) would increase the fraction of statistically significant $t$-scores by 7.8 percentage points with a standard error of only 0.6 percentage points.
I conclude that RCTs published in top economics journals are on average relatively insensitive to counterfactual sample size increases. This can only happen if most trials are either studying very large effects relative to their status quo sample sizes (and are fully powered) or are investigating effects so small that even an experiment twice as large would not easily detect them. Power calculations---however imperfect---appear to be giving researchers enough information to know when power is cheap to increase and when it is not. Since there is no clear way to check whether a single sample size was “well chosen" ex post, this new rigorous test of the aggregate downstream adequacy of power calculations was needed. The implication is that funders and researchers should generally consider devoting more resources to running more experiments or adding treatment arms rather than raising sample size standards.
This paper proposes an estimator consistent for the fraction of $t$-scores that would have been statistically significant had every experiment in a given population had its sample size counterfactually increased by a chosen factor $c^2>1$. Unlike existing work, no assumptions were imposed on the distribution of true intervention treatment effects---point masses and arbitrary densities are both allowed. The lack of such assumptions is important in theory and affects conclusions in practice.\footnote{In Table (ref) the Z-curve method of zcurve2 (which relies on stronger assumptions) finds 42% more sensitivity than the method proposed by this paper.} The proposed estimator is asymptotically normal and robust to simple forms of publication bias. A key technical contribution was to prevent uncertainty about publication bias from magnifying the smoothing error of the deconvolution step.
The method is useful and can inform funding and design decisions. For example, an empirical application finds that the power of randomized trials in economics would only increase by 7.2 percentage points on average if every sample size had been doubled. I argue that this number is small by comparing it to three benchmarks. I conclude that---despite requiring guesswork about unknown parameters---power calculations appear to leave little easy power on the table in practice. Improvements to measurement quality, compliance, and outcome choice may therefore offer lower-hanging fruit McKenzie. This suggests that funders looking to improve power should sponsor higher-quality experiments and not rely only to larger samples. Future empirical work could apply the proposed method to specific subsets of experiments, e.g. those funded by different initiatives, in order to determine exactly which sorts of experiments should be made larger.