Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
133,248 characters · 22 sections · 48 citation commands
Quasi-Bayesian Hierarchical Models
\noindentKeywords: Quasi-Bayesian inference; Generalized Bayes; Hierarchical shrinkage; GMM; Weak identification.\\ JEL codes: C11, C14, C18, C23, C26, C51.
Applied economists increasingly draw evidence from collections of related settings rather than from a single data set: for example, a treatment may be evaluated in several countries or experimental sites, or a structural parameter may be estimated separately across markets, firms, villages, or policy environments. In such cases, the empirical question is not only whether an effect is present in one sample; it is also how effects or structural parameters vary across contexts, and how evidence from one context should inform estimates in another. This is the external-validity problem: when are studies similar enough that information should be shared, and how should that sharing be done?\footnote{See, for example, Angrist2010 on the credibility revolution in internal validity, and Meager2019,Vivalt2020,Slough2022,Egami2023 for recent work on external-validity syntheses. Recent Bayesian work on prior-study information with uncertain external validity includes FinanPouzo2026,You2026.}
Bayesian Hierarchical Models (BHMs) provide a natural answer when each study can be described by a likelihood. To use them, a researcher specifies a sampling model within each group and an upper-level distribution for the group-specific parameters, and the posterior then combines the direct evidence from each group with information estimated from the collection of groups. This produces shrinkage: noisy group-level estimates are pulled toward values that are more plausible in light of the broader evidence. Hierarchical models have therefore become a standard tool for combining related treatment-effect estimates, especially when the within-study estimates are simple and approximately Gaussian; see, for example, GelmanHill2007 and Meager2019.
Many econometric applications, however, are not naturally likelihood-based. Researchers often estimate the parameter of interest using instrumental variables, GMM, minimum distance, simulated moments, nonlinear least squares, or a structural objective chosen for its identifying content rather than for its interpretation as a full sampling model. In these settings, imposing a parametric likelihood only to obtain hierarchical pooling can be unattractive. It may add distributional assumptions that are not part of the original analysis, and it may move the within-group estimand away from the parameter defined by the original econometric criterion.
This paper develops Quasi-Bayesian Hierarchical Models (QBHMs), a framework for hierarchical pooling when the within-group analysis is based on a GMM objective function, rather than a likelihood. Instead of imposing a likelihood, the construction starts from the objective function that the researcher would already use in each group. For group \(j\), let \(\phi_j\) denote the observation-level moment, residual, score, or minimum-distance map chosen by the researcher, and let \(q_{2,j,n_j}(x_j,\alpha_j)=-\|G_{j,n_j}(x_j,\alpha_j)\|_{W_{j,n_j}}^2/2\), with \(G_{j,n_j}(x_j,\alpha_j)=n_j^{-1/2}\sum_i\phi_j(X_{ji},\alpha_j)\), be the objective function the researcher would use if the group were analyzed on its own. Let \(q_{1,j}(\alpha_j,\theta;z_j)\) describe the pooling relation between the economically comparable component\footnote{Comparability means that the researcher has selected component functions or transformations \(r_j(\alpha_j)\) that have the same substantive interpretation across groups.} of \(\alpha_j\) and a hierarchical parameter \(\theta\); that is, \(q_{1,j}\) assigns higher values to configurations in which the selected group component is more compatible with the hierarchical value \(\theta\). A QBHM forms the quasi-posterior, the normalized exponential weighting of this composite objective, \[ \Pi^{\mathrm{full}}_{n,\lambda}(d\alpha\,d\theta\mid x,z) \propto \exp\!\left[ \sum_{j=1}^J q_{2,j,n_j}(x_j,\alpha_j) + \lambda\sum_{j=1}^J q_{1,j}(\alpha_j,\theta;z_j) \right] \pi_\theta(d\theta)\prod_{j=1}^J \pi_j(d\alpha_j), \] where $\propto$ denotes that the left-hand and right-hand sides differ by a constant which does not depend on $\alpha$ or $\theta$. The scalar \(\lambda\ge0\) is the pooling strength.\footnote{We normalize the lower-level quasi-Bayes temperature, meaning the scalar multiplier on the lower-level objective, to one. Fixed temperatures or learning rates are common in generalized-Bayes updates and can matter for quasi-posterior dispersion, finite-sample robustness, and calibration HolmesWalker2017,LyddonHolmesWalker2019. Under the strong-identification Laplace approximation, however, a fixed temperature does not change the first-order location of the Laplace-type estimator (LTE), and in the present normalization the substantively reported scale parameter is \(\lambda\), the relative weight on pooling.} Throughout the paper, the hierarchy refers to this upper-level pooling relation: the combination of \(q_1\), \(\pi_\theta\), and \(\lambda\) that specifies which study-level components should be similar and how strongly they are pulled together. It is a researcher-specified pooling device, not a requirement that the studies be generated from a random-effects population. At \(\lambda=0\), the marginal quasi-posteriors for each component $\alpha_j$ are the same as would be obtained by separate Laplace-Type Estimation ChernozhukovHong2003. Furthermore, when the $q_{2,j,n_j}$ functions are log likelihoods, \(q_{1,j}\) is a conditional log density, and \(\lambda=1\), this replicates a classical Bayesian Hierarchical Model. The lower level therefore remains an econometric objective function, while the upper level supplies hierarchical shrinkage. Our analysis in this paper considers the properties of the quasi-posterior mean, which we will denote the QBHM estimator; we treat its components as point estimators of the study-level estimands $\alpha_1, ..., \alpha_J$. Our theory considers the case where $J$ is a fixed number (rather than growing with $n$), as the number of studies is typically very small compared to the sample size. One consequence of this is that we do not consider estimation of $\theta$ (such as the so-called `average Average Treatment Effect'), as this is not generally identified unless $J \rightarrow \infty$.
The paper makes four main contributions. First, we introduce and highlight the flexibility of the QBHM framework. The group objective functions, instrument sets, controls, fixed effects, auxiliary statistics, weights, and dimensions may differ, as long as the pooled components are economically comparable. As such, one study can be estimated by IV with its own instruments, another by a fixed-effects residual objective function, another by a score objective function from a quasi-likelihood, and another by minimum distance on auxiliary moments. The hierarchy only links the components the researcher chooses to pool; these may be site-level treatment effects, IV returns to capital, production elasticities, switching thresholds, or other estimands with the same economic interpretation across studies.
Second, we give fixed-\(J\) theory under grouped-GMM conditions that are directly checkable in smooth applications. We formalize strong and weak identification by using differentiability in quadratic mean (DQM) local paths, defined below as \(O(n_j^{-1/2})\)-local perturbations of the group data law, following the weak-identification perspective in kaji2021weakidentification and the weak-GMM limit experiment in AM2022. Strong studies have the same asymptotic distribution as their unpooled GMM or LTE analogues, as in a standard IV design with a first stage bounded away from zero. Weak studies, such as IV designs with first stages local to zero or nonlinear GMM problems with weak curvature, remain in a weak-GMM limit experiment, meaning a limit in which the whole sample-moment objective remains a random function over the weak parameter space. In this limit, the hierarchy changes the \(\lambda\)-indexed prior over weak values, which we call the weak-limit prior path, and therefore remains in the asymptotic distribution. The appendix gives the corresponding profiled treatment of studies containing both strongly and weakly identified parameters.
Third, we characterize when our procedure can have lower asymptotic MSE, compared to estimating studies individually by Laplace-Type methods or GMM. Intuitively, pooling creates a bias--variance tradeoff, which can lead to lower MSE when the variance reduction dominates the increase in bias, in the same broad shrinkage-risk spirit as JamesStein1961. A local nonlinear weak-GMM result and a scalar weak-IV example make this condition explicit in two salient cases.
Finally, we give practical guidance for reporting the pooling strength $\lambda$ and for statistical inference. We recommend reporting sensitivity paths over \(\lambda\), rather than demonstrating results over only one value of $\lambda$, though a form of cross validation (CV) provides a heuristic for a good value. Drawing analogies to AM2016conditional, we show how one might construct weak-identification-robust confidence sets for hypothesis testing, noting that these do not agree with using the quasi-posterior itself.
\paragraph{Existing Literature.} The paper is related to several strands of work. First, generalized Bayes and moment-based Bayesian procedures replace a full likelihood with a loss, moment condition, or empirical-likelihood object in the update BissiriHolmesWalker2016,ChernozhukovHong2003,Schennach2005BayesianETEL,ChibShinSimoni2018MomentConditions. Recent work develops related quasi-Bayesian procedures for moment misspecification and conditional moment restriction settings ChernozhukovHansenKongWang2025PlausibleGMM,Kankanala2025GeneralizedBayesCMR. The present paper differs because we consider \(J\) to be fixed, and moreover study the shrinkage induced by the hierarchy, rather than misspecification robustness, frequentist coverage, or contraction for a single quasi-posterior. The closest conceptual comparison to our work is the quasi-Bayesian grouped-panel framework of Huang2023GroupedPanels, which also combines flexible loss functions and priors for grouped data, but the object of interest is different. In our setting, the groups are observed econometric problems rather than latent panel clusters, the lower-level objective functions may differ across groups, and the theory focuses on fixed-\(J\) hierarchical shrinkage and weak-GMM MSE rather than posterior contraction for latent group assignments.
Hierarchical Bayes, empirical Bayes, and shrinkage show why pooling can reduce quadratic risk in normal means, multi-study, and labor-economics settings JamesStein1961,Meager2019,Walters2024EmpiricalBayes. QBHM differs because the lower level need not be a common likelihood, a common normal approximation to a reduced-form summary, or a many-groups population-distribution problem. Penalized or regularized GMM is also related to our work, since a QBHM mode maximizes the original grouped objective function plus an induced penalty. The object analyzed in this paper, however, is the quasi-posterior mean \(\tilde\alpha^{\mathrm{full}}_{n,\lambda}\). The mode and mean are approximately the same in the strong identification case, by the generalized Bernstein von-Mises theorem in ChernozhukovHong2003. However, this equivalence need not hold under weak identification, as shown in AM2022, since the posterior is no longer asymptotically Gaussian; in general, it can be skewed, with the mean no longer equal to the mode. Relatedly, AndrewsMikusheva2023Inadmissible show that GMM estimators can be inadmissible under weak identification, while quasi-Bayes posterior means and bootstrap-aggregated (bagged) GMM have stronger continuity properties. QBHM adds a cross-group hierarchy to this weak-GMM decision environment.
Finally, weak-instrument, semiparametric weak-identification, and weak-GMM work show that conventional first-order approximations can fail when the objective function is weakly curved or has only finite local drift StaigerStock1997,StockYogo2005,kaji2021weakidentification. Kaji's formulation explains weak identification through local paths that do not regularize the estimand at the root-\(n\) scale, and the Andrews--Mikusheva weak-GMM experiment supplies the decision-theoretic environment that QBHM uses below for point-estimation MSE. For inference, the conditional-testing approach builds on the weak-instrument and functional-nuisance conditioning literatures Moreira2003,AndrewsMoreiraStock2006,AM2016conditional. As in that literature, quasi-posterior intervals summarize the shrinkage estimator, while weak-identification-robust confidence sets should still be constructed by conditioning.
The paper proceeds as follows. Section (ref) gives the normal--normal example that introduces the MSE tradeoff and the strong-versus-weak information scale. Section (ref) defines the general grouped QBHM and gives the fixed-\(J\) theory, including the strong-study approximation, the weak-study weak-GMM reduction, and the Bayes-rule interpretation under squared loss of the rule induced by the hierarchy. Section (ref) states the pointwise MSE comparison, a local sufficient condition, and a scalar weak-IV example. Section (ref) discusses sensitivity paths and optional CV-selected reference choices for \(\lambda\), together with weak-identification-robust confidence sets. Sections (ref) and (ref) present the Monte Carlo evidence and empirical illustration. Section (ref) concludes.
To fix intuition, we begin with an exact Bayesian hierarchical model. The rest of the paper does not impose these distributional assumptions, but the normal--normal case is useful because it gives closed-form shrinkage coefficients and makes the role of identification strength transparent. It is also close to models used in applied hierarchical analyses, and so the results here may be of independent interest.\footnote{For example, Meager2019 considers a model with normal distributions in the main stages, apart from the prior on the heterogeneity parameter. If that heterogeneity parameter were fixed, the model would fall within the case considered here.} Suppose that there are \(J\) observed groups, and that group \(j\) is summarized by a scalar estimate \(\Delta_j\) of a group-specific estimand \(\alpha_j\), such as a site-level treatment effect or IV slope. Specifying a Bayesian hierarchical model requires us to specify our beliefs over the data in three parts. In this case, the lower level is \[ \Delta_j\mid \alpha_j \sim N\!\left(\alpha_j,\frac{\sigma_{\Delta,j}^2}{\mathcal I_{j,n}}\right), \qquad \sigma_{\Delta,j}^2>0,\quad \mathcal I_{j,n}>0 . \] The hierarchical level is \[ \alpha_j\mid \theta \sim N(\theta,\tau_0^2), \qquad \tau_0^2>0 , \] and the prior level is \[ \theta\sim N(g_0,s_0^2), \qquad s_0^2>0 . \] Here \(\theta\) is the common mean toward which the group-specific estimands are pooled, \(\tau_0^2\) controls prior cross-group dispersion around that mean, and \(\mathcal I_{j,n}/\sigma_{\Delta,j}^2\) is the information in the group-\(j\) summary. We use \(\mathcal I_{j,n}\) to model identification strength in this setup: strong identification corresponds to \(\mathcal I_{j,n}=n_j\to\infty\), so the group summary becomes increasingly precise, while weak identification corresponds to \(\mathcal I_{j,n}\) remaining bounded, so the effective information in \(\Delta_j\) does not grow with \(n_j\).
The joint quasi-posterior is proportional to \[ \exp\left[ -\frac12\sum_{j=1}^J\frac{\mathcal I_{j,n}}{\sigma_{\Delta,j}^2}(\Delta_j-\alpha_j)^2 -\frac{\lambda}{2}\sum_{j=1}^J\frac{(\alpha_j-\theta)^2}{\tau_0^2} -\frac12\frac{(\theta-g_0)^2}{s_0^2} \right]. \] When \(\lambda=1\), this is the ordinary posterior under the normal hierarchical model above; when \(\lambda=0\), the groups are estimated separately; and, more generally, \(\lambda/\tau_0^2\) is the effective pooling precision. Increasing \(\lambda\), or decreasing \(\tau_0^2\), strengthens the pull of each \(\alpha_j\) toward the common mean. The same display can also be read as a quadratic QBHM objective. The lower normal density contributes \[ q_{2,j,n_j}^{\mathrm{N}}(\Delta_j,\alpha_j) = -\frac12\frac{\mathcal I_{j,n}}{\sigma_{\Delta,j}^2} (\Delta_j-\alpha_j)^2, \] the hierarchy contributes \[ q_{1,j}^{\mathrm{N}}(\alpha_j,\theta) = -\frac12\frac{(\alpha_j-\theta)^2}{\tau_0^2}, \] and the prior contributes the final quadratic term in \(\theta\). The exponent in the quasi-posterior is therefore the composite objective \[ \sum_{j=1}^J q_{2,j,n_j}^{\mathrm{N}}(\Delta_j,\alpha_j) + \lambda\sum_{j=1}^J q_{1,j}^{\mathrm{N}}(\alpha_j,\theta) - \frac12\frac{(\theta-g_0)^2}{s_0^2}, \] Setting \(\lambda=1\) gives the ordinary normal hierarchical posterior for this benchmark specification. The convenient functional forms yield explicit expressions for the marginal quasi-posterior mean of $\alpha_j$, denoted $\tilde\alpha_{j,\lambda}$, which is the QBHM estimator that we study in this paper.
In this case, we see that the QBHM estimates are a weighted average of individual estimates and a term which combines information from all of the groups. Note that when \(\lambda=0\) we have no pooling, and consequently \(\tilde\alpha_{j,0}=\Delta_j\). For fixed \(\lambda>0\), a group with weaker identification, meaning a larger sampling variance \(\sigma_{\Delta,j}^2/\mathcal I_{j,n}\), has a smaller \(c_{j,n}(\lambda)\) and so receives more shrinkage toward \(\tilde\theta_\lambda\). Conversely, when the group summary is precise, \(c_{j,n}(\lambda)\) is close to one and the pooled estimate remains close to \(\Delta_j\).
Proposition (ref) contains two sources of shrinkage: each \(\tilde\alpha_{j,\lambda}\) is shrunk toward the estimated common mean \(\tilde\theta_\lambda\), and \(\tilde\theta_\lambda\) itself is a compromise between the prior mean \(g_0\) and the group summaries. When \(s_0^2\) is small, the prior pulls the common mean toward \(g_0\), but when \(s_0^2\) is large, this direct prior pull is weak. As the prior gets wider and wider, as in (ref), the common mean instead converges to an empirical average of the group summaries, with weights determined by the shrinkage coefficients. Weakly identified groups are therefore pulled more toward the common mean, but they also receive less weight in forming that common mean. We can see from this that Bayesian Hierarchical models naturally provide an identification strength-adaptive way to pool data together from multiple studies.
One common concern when using these methods is the influence that the prior distribution has over the results. As discussed, we treat the mean of the quasi-posterior as a point estimator, and analyse its frequentist properties, rather than consider QBHMs as Bayesian objects. The standard solution to the choice of priors is twofold. First, as shown in our main results, prior choice is asymptotically negligible for the estimator, when working under strong identification. The prior does matter in small samples and under weak identification - as highlighted above - but this can be viewed as strength of the method, as it allows the incorporation of extra information. This is especially important in the case of weakly identified parameters, where the data themselves do not pin down the estimand consistently. Second, a researcher who wishes to de-emphasize the effect of priors can set them to be very wide. This is commonly done in applied work (e.g. by Meager2019) and, as seen from diffuse-prior limit above, makes the shrinkage effect of the priors become asymptotically irrelevant. All of these features carry through to the QBHM case. As such, our method provides the flexibility to use prior information if needed, but with no real obligation to do so.
The next result records the consequences of identification strength. It shows that fixed pooling has no effect on the asymptotic distribution when group \(j\) is strongly identified, but it does have an effect when the identification is weak.
Note from (ref) that, in the diffuse-prior limit and for \(\lambda>0\), the shrinkage target is not necessarily the true common mean. Instead, it is the weighted empirical mean \(\tilde\theta_\lambda^{\mathrm{diff}}\) of the group summaries \(\Delta_k\). Thus the pooled estimate is pulled toward an empirical center that includes group \(j\)'s own summary, rather than toward an oracle or leave-one-out target. Pooling is therefore most likely to improve MSE when the unpooled summary is noisy and this empirical center is well aligned with the group-specific estimand. On the other hand, it is likely to harm MSE when the summary is already precise or the center is poorly aligned with the group.
The next two results characterize when pooling can improve a form of mean squared error. The first fixes \(\theta\) in a simplified known-center case to isolate the scalar signal-to-noise tradeoff behind the shrinkage coefficient and to evaluate integrated MSE. That is, it asks whether pooling can lower mean squared error before conditioning on the realized group-specific effect.\footnote{This can be viewed as ex ante MSE, or equivalently as average MSE under the hierarchical distribution. It differs from the pointwise MSE below, which conditions on the realized value of \(\alpha_{j0}\) and averages only over the sampling error.} For this proposition only, suppose that, conditional on known \(\theta\), \(\alpha_j=\theta+u_j\) with \(u_j\sim N(0,\tau^2)\), and that \(\Delta_j=\alpha_j+\varepsilon_j\) with \(\varepsilon_j\sim N(0,\sigma_{\Delta,j}^2/\mathcal I_{j,n})\) independent of \(u_j\); write \(v_j=\sigma_{\Delta,j}^2/\mathcal I_{j,n}\).
The integrated-risk-optimal level of pooling has an intuitive interpretation. The parameter \(\tau^2\) measures how dispersed the true group effects are around the center, while \(v_j\) measures how noisy the group summary is. When \(v_j\) is small relative to \(\tau^2\), the summary \(\Delta_j\) is precise, the center is weakly predictive of \(\alpha_j\), or both, so \(\eta_j^\star\) is close to one and the integrated-risk-optimal rule puts little weight on the center. When \(v_j\) is large relative to \(\tau^2\), the summary is noisy, the group effects are expected to lie close to the center, or both, so \(\eta_j^\star\) is smaller and the optimal rule pools more aggressively.
The last statement of Proposition (ref) calibrates $c_{j,n}(\lambda)$ against this integrated-risk benchmark. Since \(c_{j,n}(\lambda)=1/(1+\lambda v_j/\tau_0^2)\), the QBHM puts less weight on the unpooled summary when \(v_j\) is large, when \(\lambda\) is large, or when \(\tau_0^2\) is small. Weak information, stronger imposed pooling, and tighter hierarchical dispersion all push in the same direction. The coefficient is integrated-risk optimal in this case only when the effective pooling precision \(\lambda/\tau_0^2\) matches the random-effects precision \(1/\tau^2\). If the imposed pooling precision is larger than \(1/\tau^2\), the rule shrinks more than the optimum; if it is smaller, the rule shrinks less.
In contrast with the last result, the next corollary conditions on the realized group-specific estimand and asks when shrinkage toward a fixed center improves on the unpooled summary. This distinction matters because pooling can be attractive on average under a well-centered random-effects distribution, while still increasing MSE for a group whose realized estimand is far from the center.
The pointwise condition separates the variance gain, \(v_j\), from the bias cost, \((\theta-\alpha_{j0})^2\), the squared distance between the shrinkage target and the realized estimand. Pooling is therefore most likely to help when \(v_j\) is large and \(\theta\) is close to \(\alpha_{j0}\): the unpooled estimate is noisy, and the bias from moving toward the center is small. Pooling is most likely to hurt when \(v_j\) is small or when \(\theta\) is far from \(\alpha_{j0}\): the unpooled estimate is already precise, or the shrinkage target is badly aligned. Full shrinkage to the center corresponds to \(\eta=0\), in which case pooling improves pointwise MSE exactly when \((\theta-\alpha_{j0})^2<v_j\). As \(\eta\) approaches one, the condition becomes easier to satisfy because the amount of shrinkage is small, but the possible MSE gain also becomes small.
Proposition (ref) and Corollary (ref) describe the same shrinkage tradeoff from two perspectives. Proposition (ref) treats the center as known and evaluates MSE after averaging over a normal random-effects distribution. Corollary (ref) fixes the realized estimand and evaluates MSE pointwise. Together, they show why pooling can help in weak-information settings: it reduces variance when the unpooled summary is noisy, but it can introduce bias when the center is not well aligned with the group-specific estimand.
The full QBHM in Proposition (ref) is complicated not only because the objective functions need not correspond to Gaussian log likelihoods, but also because the hierarchy may use a general prior and need not have a known center. The next section states the grouped GMM QBHM directly, with separate first-order treatments of strong and weak study blocks.
We now state the regularity conditions needed for our main theory. Section (ref) defines strong and weak identification, Section (ref) gives the result for strongly identified studies, and Section (ref) gives the result for weakly identified studies; throughout this section, “group” and “study” are used interchangeably. For the main text, identification strength is the same for all components in a study: the reported vector for a study is either strongly identified or weakly identified. Appendix (ref) gives the extension that allows mixed identification status within a study. Throughout, stochastic processes indexed by compact sets are taken to have Borel measurable versions in the corresponding sup-norm spaces.\footnote{This is a measurability convention for process-level convergence statements: if an empirical-process construction gives only asymptotic measurability, weak convergence and probability statements are understood in the usual outer-probability sense.} For a symmetric positive semidefinite matrix \(A\), write \(\|x\|_A^2:=x'Ax\).
Consider \(J\ge1\) independent groups, where group \(j\) has data \(x_j=(x_{j1},\ldots,x_{j n_j})\), fixed group-level covariates \(z_j\) when such covariates enter the pooling relation, and parameter \(\alpha_j\in\mathcal A_j\subset\mathbb{R}^{d_j}\). Stack \(\alpha=(\alpha_1',\ldots,\alpha_J')'\in\mathcal A:=\prod_{j=1}^J\mathcal A_j\), allowing the dimensions \(d_j\) to differ across groups. The pooling relation may therefore be applied only to economically comparable components or transformations of them, denoted \(r_j(\alpha_j)\), leaving study-specific nuisance components unpooled; for example, \(r_j\) may select a treatment effect, an IV slope, a production elasticity, or a switching threshold from a larger study-specific parameter vector. The hierarchical parameter \(\theta\in\Theta\subset\mathbb{R}^{d_\theta}\) indexes the function that creates pooling. Let \(n_j\) denote the sample size of group \(j\).
At any law \(P_j\) at which the group-\(j\) estimand is identified, write \(\alpha_j(P_j)\) for the estimand.\footnote{The theory also allows pseudo-true estimands. For example, under fixed-weight GMM, \(\alpha_j(P_j)\) may be any maximizer of the population objective \(a\mapsto-\|\mathbb E_{P_j}\phi_j(X_{ji},a)\|_{W_j}^2/2\).} The moment condition defining that estimand is
where \(\phi_j:\mathcal X_j\times\mathcal A_j\to\mathbb{R}^{k_j}\) is chosen by the researcher. When this restriction has multiple solutions, the notation refers to the solution, selection, or pseudo-true value specified by the study-level criterion. It is important that we explicitly consider a map $P_j \mapsto \alpha_j(P_j)$, rather than only look at the moment condition under one `true' data distribution, as this is needed to study weak identification. We follow the formalization in kaji2021weakidentification below, which nests several common definitions. The data need not be generated from a random-effects model; while a plausible random-effects description may help guide the choice of the pooling term, the study-level estimand is defined through the moment condition rather than through a random-effects law. Let \(\hat g_j(a):=n_j^{-1}\sum_i\phi_j(X_{ji},a)\) and \(G_{j,n_j}(x_j,a):=\sqrt{n_j}\,\hat g_j(a)\). With positive definite weight \(W_{j,n_j}\), the within-group objective function is
This display is the fixed-weight GMM objective in maximization form. With continuously updated GMM, the weight depends on the candidate value \(a\); for example, \(W_{j,n_j}(a)=\widehat\Sigma_{j,n_j}(a,a)^{-1}\), where \(\widehat\Sigma_{j,n_j}(a,\tilde a)\) estimates the covariance between \(G_{j,n_j}(X_j,a)\) and \(G_{j,n_j}(X_j,\tilde a)\). The weak-study subsection uses a generic estimated weight process for the lower-level objective, though in the continuously updated case, that process is the diagonal inverse-covariance weight. Write \(\bar q_{2,j,n_j}=n_j^{-1}q_{2,j,n_j}\), and let \(M_j\) denote the population criterion that is the uniform limit of \(\bar q_{2,j,n_j}\) for strongly identified studies. In fixed-weight GMM, if \(\mu_j(a):=\mathbb E_{P_{j0}}\phi_j(X_{ji},a)\) and \(W_{j,n_j}\overset{p}{\to} W_j\), then \[ \bar q_{2,j,n_j}(x_j,a) = -\frac12\hat g_j(a)'W_{j,n_j}\hat g_j(a), \qquad M_j(a)=-\frac12\|\mu_j(a)\|_{W_j}^2 . \] As discussed in earlier sections, the pooling term is a second objective function \(q_{1,j}(\alpha_j,\theta;z_j)\). In an ordinary likelihood-based hierarchy, it would be the log density of \(\alpha_j\) given \((\theta,z_j)\). In a QBHM, however, it may instead be any continuous pooling objective, such as a smooth penalty that rewards similarity in \(r_j(\alpha_j)\). For example, a Gaussian common-mean term for comparable components is \[ q_{1,j}(\alpha_j,\theta) = -\frac12(r_j(\alpha_j)-\theta)'\Omega_{0j}^{-1}(r_j(\alpha_j)-\theta), \] up to constants. Set \[ q_{2,n}(\alpha;x):=\sum_{j=1}^J q_{2,j,n_j}(x_j,\alpha_j), \qquad q_1(\alpha,\theta;z):=\sum_{j=1}^J q_{1,j}(\alpha_j,\theta;z_j), \] with \(z=(z_1,\ldots,z_J)\). The observations \(X_{ji}\) enter the group objective; the fixed covariates \(z_j\) enter the pooling relation. For a pooling-strength parameter \(\lambda\ge0\), define
and let \(\pi(d\alpha\,d\theta)=\pi_\theta(d\theta)\prod_{j=1}^J\pi_j(d\alpha_j)\) be a proper product prior. The QBHM quasi-posterior is
Recall that if \(q_{2,j,n_j}\) is a group log likelihood, \(q_{1,j}\) is a conditional log density, and \(\lambda=1\), this display is ordinary hierarchical Bayes. Otherwise, the lower level remains the econometric objective function in (ref). At \(\lambda=0\), the marginal law of \(\alpha\) is the product of the unpooled quasi-posteriors, \[
\] The parameter \(\theta\) then has no effect on the marginal estimates of the study-level estimands. Under squared loss, the QBHM estimator is the marginal quasi-posterior mean
Throughout the paper, a tilde denotes a marginal quasi-posterior mean and a circumflex denotes an objective-function maximizer. The choice of \(\lambda\) determines the position on the bias--variance tradeoff. Section (ref) therefore recommends reporting estimates over a sensitivity path in \(\lambda\), rather than only at one value.
The first formal condition fixes the sampling environment, compact supports, continuity needed for the quasi-posterior, and base-prior regularity.
In applications, compactness can be interpreted as a sufficiently wide truncation chosen before the asymptotic approximation is applied. Assumption (ref) imposes independent sampling units within groups and independent samples across groups, but it does not impose a common likelihood, a common moment dimension, a common instrument set, or common supports across groups. A collection of studies with different instruments, controls, fixed effects, or auxiliary statistics can therefore be pooled when the components linked by \(q_1\) have a common interpretation. The notation is abstract, but the requirement is the standard GMM one: each study supplies a moment objective for its own estimand, and the hierarchy only adds a cross-study relation for selected comparable components. The compactness, continuity, and density conditions are used to make the quasi-posterior proper and to give uniform bounds in the reductions below.
This section defines what it means for a reported study block to be strongly or weakly identified. Recall that the data law in group \(j\) is denoted \(P_{j0}\). In the strongly identified and point-identified case, we write \(\alpha_{j0}:=\alpha_j(P_{j0})\), so the estimand is well defined at this fixed law. Weak identification requires a slightly different formalization. We treat \(P_{j0}\) as a reference law at which the estimand may be only set identified, and study sequences of nearby laws \(P_{j,n_j,f_j}\) that converge to \(P_{j0}\) at the \(O(n_j^{-1/2})\) scale. Along these sequences, the estimand is point identified, so the local path selects a particular nearby value from the set of values consistent with \(P_{j0}\). Equivalently, when the sample size in group \(j\) is \(n_j\), we view the data as drawn from \(P_{j,n_j,f_j}\) rather than exactly from \(P_{j0}\). This formalization, originally due to kaji2021weakidentification, covers several familiar forms of weak identification. For example in weak IV, the correlation between the instrument and the endogenous regressor converges to $0$ as the sample size increases, so the sequence \((P_{j,n_j,f_j})_{n_j}\) consists of laws under which the first stage becomes weak at the same rate as sampling noise\footnote{Rank-deficient GMM and local-to-singular moment problems fit the same description. When regular and weak coordinates appear inside the same study, Appendix (ref) profiles the regular coordinates and leaves a reduced moment process with Gaussian variation and finite local drift.}. The direction \(f_j\) indexes which local departure from \(P_{j0}\) is being considered: in the weak-IV example, it determines the local first-stage strength, while in a rank-deficient GMM problem it determines the local drift that selects among otherwise observationally similar parameter values.
Since groups are independent, local sequences can be specified group by group: for \(f=(f_1,\ldots,f_J)\), \(P_{n,f}=\bigotimes_{j=1}^J P_{j,n_j,f_j}^{n_j}\). Thus local identification is defined within each group and then combined across groups by independence. The relevant notion of convergence is differentiability in quadratic mean, or DQM. After choosing a common dominating measure \(\mu_j\),\footnote{For example, Lebesgue measure or any measure dominating the local path.} let \(p_{j,n_j,f_j}\) and \(p_{j0}\) denote the densities of \(P_{j,n_j,f_j}\) and \(P_{j0}\).
Write \(L_2^0(P_{j0})=\{f\in L_2(P_{j0}):\mathbb E_{P_{j0}}f=0\}\). This formulation is useful because it separates the original structural model from the local weak experiment. Let \(\mathcal P_{j,\alpha}\) denote the part of the group-\(j\) model on which the reported estimand \(\alpha_j(P_j)\) is point identified. In a weak-identification sequence, the reference law \(P_{j0}\) may be an identification-failure law, so \(\alpha_j(P_{j0})\) need not be a single well-defined structural value. We therefore look at DQM local paths that approach \(P_{j0}\) while remaining in \(\mathcal P_{j,\alpha}\). A score \(f_j\) is called pertinent for the reported estimand if at least one such DQM path induces score \(f_j\) and all such paths with the same score have the same limiting estimand value. Write \(\dot{\mathcal P}_{j0,\alpha}\) for the resulting pertinent tangent cone, and write \[ \alpha^W_{j0}(f_j) := \lim_{n_j\to\infty}\alpha_j(P_{j,n_j,f_j}) \] for the common limit when it exists\footnote{The cone language matters: multiplying the local score by a positive constant changes the speed along the same local ray but not the limiting weak value. Thus the weak estimand is allowed to depend on the direction of approach to the failure point, but not linearly on the magnitude of the perturbation.}. We can now define identification strength as follows:
Whether a reported estimand is strongly or weakly identified is a property of the map \(P\mapsto\alpha_j(P)\), not of a particular estimator. Strong identification means that the estimand has a first-order linear derivative with respect to local changes in the law. Weak identification means instead that the estimand has a well-defined limit along each pertinent local direction, but this limit is a nonlinear direction-only object. That direction-only dependence is the source of the non-Gaussian weak-GMM limit below.
For weak GMM, there is one additional reduction before the process limit is written down. We keep \(\mathcal A_j\) for the researcher's original candidate set and define the reference-law population moment and zero set by \[ \mu_{j0}(a)=\mathbb E_{P_{j0}}\phi_j(X_{ji},a), \qquad \mathcal Z_{j0}=\{a\in\mathcal A_j:\mu_{j0}(a)=0\}. \] The weak-GMM limit is indexed by a retained weak space: a compact set \(\mathcal W_j\subseteq\mathcal Z_{j0}\) that contains the weak limits \(\alpha^W_{j0}(f_j)\) for the local directions under consideration. The limit below is indexed by \(\mathcal W_j\), not by all of \(\mathcal A_j\). Values in \(\mathcal A_j\setminus\mathcal Z_{j0}\) do not have a finite weak-process limit, since \(g_{j,n}(a)=\sqrt{n_j}\mu_{j0}(a)+O_p(1)\) under the reference law. They are therefore regularized, profiled out, or asymptotically discarded, rather than included in the weak experiment. If \(\mathcal Z_{j0}\) is lower-dimensional, the retained-space base measure introduced below is understood to be a fixed coordinate, Hausdorff, or other appropriate measure on \(\mathcal W_j\).
When each reported estimand is scalar, this condition only requires classifying each study as strong or weak. The condition is imposed only for the main fixed-\(J\) theory; Appendix (ref) gives the profiled extension for a study containing both strongly and weakly identified components.
We now give primitive conditions for the strong-study results. The following condition gives a locally quadratic population objective and enough sample regularity for the corresponding lower-level quasi-posterior to concentrate at the usual \(O(n_j^{-1/2})\) rate.
These high-level objective conditions for smooth finite-dimensional GMM impose the local criterion behavior needed for the quasi-posterior reduction once a study is classified as strong, without requiring the objective to be a likelihood. Assumption (ref) is checked by verifying that the population criterion has a single interior maximizer, has nonsingular curvature there, and is uniformly approximated by the sample criterion and its first two derivatives near that point under the triangular-array law being studied. For a smooth finite-dimensional GMM criterion, this follows from uniform laws of large numbers for the moments and their derivatives, convergence of the weighting matrix when the criterion uses one, and a central limit theorem for the score at \(\alpha_{j0}\). The Hessian condition is often implied by full rank of the moment Jacobian under correct specification; under misspecification it can be checked directly from the population criterion. Under Assumption (ref), with probability approaching one there exists \(\widehat\alpha_j^{\mathrm{un}}\) such that \[ \widehat\alpha_j^{\mathrm{un}} \in \arg\max_{a\in\mathcal U_j(\delta_j)} q_{2,j,n_j}(X_j,a), \qquad \sqrt{n_j}\bigl(\widehat\alpha_j^{\mathrm{un}}-\alpha_{j0}\bigr) =O_p(1). \] Here \(\widehat\alpha_j^{\mathrm{un}}\) is the standard GMM estimator, and Theorem (ref) is stated as a quasi-posterior law result in the local coordinate centered at this maximizer. In the result below, fix \(j\in\mathcal J_S\), and let \(\widehat\alpha_j^{\mathrm{un}}\) be a measurable local maximizer over \(\mathcal U_j(\delta_j)\) satisfying \(\sqrt{n_j}(\widehat\alpha_j^{\mathrm{un}}-\alpha_{j0})=O_p(1)\). Let \(P^S_{j,n,\lambda}\) be the marginal law under the full QBHM quasi-posterior of \(h_j=\sqrt{n_j}(\alpha_j-\widehat\alpha_j^{\mathrm{un}})\), and let \(P^{S,0}_{j,n}\) be the law of the same \(h_j\) under the unpooled quasi-posterior \(\Pi^{\mathrm{un}}_{j,n}\).
The final result says that the marginal QBHM quasi-posterior mean is first-order equivalent to the unpooled local maximizer. More specifically, it implies that the QBHM estimator $\tilde\alpha^{\mathrm{full}}_{j,n,\lambda}$ and the usual GMM estimator $\widehat\alpha^{\mathrm{un}}_j$ have the same asymptotic distribution, no matter the value of $\lambda$ chosen. This parallels the main LTE result of ChernozhukovHong2003 - the difference here is that we focus on the quasi-posterior mean, rather than Bayes estimates defined by a more general loss function, and we show uniformity over the $\lambda$ term. For fixed pooling strengths, the theorem shows that, in a strongly identified study, the QBHM quasi-posterior mean differs from the unpooled local GMM maximizer by \(o_p(n_j^{-1/2})\), uniformly in \(\lambda\). Thus any usual first-order limit law for the unpooled maximizer carries over to the QBHM mean. In particular, uniform integrability implies that the two estimators have the same asymptotic MSE. The estimators can still differ in finite samples, and the prior can matter at that scale, but fixed pooling has no first-order effect under strong identification. Pooling can enter at first order only in the weak-study limit considered next.
For the retained weak spaces \(\mathcal W_j\) introduced in Section (ref), set \[ \mathcal W:=\prod_{j\in\mathcal J_W}\mathcal W_j, \qquad \alpha_W=(\alpha_j)_{j\in\mathcal J_W}\in\mathcal W . \] For weak study \(j\), define the local moment process \[ g_{j,n}(\alpha_j) := G_{j,n_j}(X_j,\alpha_j) = n_j^{-1/2}\sum_{i=1}^{n_j}\phi_j(X_{ji},\alpha_j), \] and stack these processes as \(g_n(\alpha_W):=\bigl(g_{j,n}(\alpha_j)'\bigr)_{j\in\mathcal J_W}'\). Let \(k_W:=\sum_{j\in\mathcal J_W}k_j\). Under the primitive conditions below, \(g_n\) converges on \(\mathcal W\) to \(g=m+\mathbb G\), and the weak-study objective converges to the grouped weak-GMM criterion \(Q\). In this notation, \(\alpha_W\) is a candidate retained weak value, \(\alpha_W^*\) is the fixed weak value selected by the local path, \(g\) is the random weak-GMM moment process, and \(m\) is its deterministic drift in the fixed weak-GMM state. Local coordinates \(h\) are used only after centering at \(\alpha_W^*\), so \(h=\alpha_W-\alpha_W^*\); when examples write a generic candidate value \(a\), it is translated into \(h\) by the same identity. The testing subsection later defines \(h_{a_0}\) as a residual process, not as this local coordinate. At this scale the hierarchy enters through the prior path \(\pi_\lambda\), so changing \(\lambda\) changes the limiting weak quasi-posterior mean even though $\phi_j$ is unchanged. The weak-study part of the QBHM lower level is
The additive term \(c_n\) may depend on the data but not on \(\alpha_W\), and therefore cancels from normalized quasi-posteriors. The display is only a reduction of the original weak-study criteria to the retained weak coordinates; it does not replace the lower-level objective by a different one.
The same weight matrix appears in the original lower-level criterion and in \(Q_n\). With fixed weights, \(\widehat W_n\) is the fixed, or converging, block-diagonal GMM weight. With continuously updated GMM, it is the estimated weight evaluated at the candidate value \(\alpha_W\). The off-diagonal covariance kernel used for feasible weak-identification-robust tests is introduced in Section (ref); it is not part of the candidate-value objective in this subsection. Mixed within-study blocks require a profiled weak experiment and are treated in Appendix (ref).
The assumption below fixes the local law, the retained weak spaces, the empirical-process limit, and the weight convergence. Let \(P^W_{n,f}:=\bigotimes_{j\in\mathcal J_W}P_{j,n_j,f_j}^{n_j}\) be the product local law for the weak studies. For the DQM score \(f_j\), define \(m_j(\alpha_j):=\mathbb E_{P_{j0}}[f_j(X_{ji})\phi_j(X_{ji},\alpha_j)]\) and \(m(\alpha_W):=(m_j(\alpha_j)')_{j\in\mathcal J_W}'\). Let \(W:\mathcal W\to\mathbb R^{k_W\times k_W}\) denote the nonrandom limit of the estimated GMM weight process.
Under Assumption (ref), \(\alpha_W^*:=(\alpha_j^*)_{j\in\mathcal J_W}\) belongs to \(\mathcal W\) and satisfies \(m(\alpha_W^*)=0\), the state restriction used later for pointwise MSE and weak-identification-robust testing. Clause (i) selects the local sequence and therefore the drift \(m\), clause (ii) records the retained-space restriction, and clauses (iii)--(iv) give the uniform central limit theorem and weight convergence needed for the weak quasi-posterior limit. For continuously updated GMM, clause (iv) is typically verified by uniform convergence and nonsingularity of the estimated covariance matrix. Under this condition, we get the following.
Proposition (ref) supplies the random criterion for the reduced weak quasi-posterior. Because \(Q\) is a process on the retained space \(\mathcal W\), rather than a local quadratic expansion around a single point, the limit is not a single point, like it would be in the strongly-identified case. Since strong studies are consistently estimated by the QBHM, the weak-study limit includes the strong studies, evaluated at their true values $\alpha_{j0}$. Write \(\mathcal A_S:=\prod_{j\in\mathcal J_S}\mathcal A_j\) and \(\alpha_{S0}:=(\alpha_{j0})_{j\in\mathcal J_S}\). For \(a_S=(a_j)_{j\in\mathcal J_S}\in\mathcal A_S\), set \(q_{1,S}(a_S,\theta;z_S):=\sum_{j\in\mathcal J_S}q_{1,j}(a_j,\theta;z_j)\) and \(q_{1,W}(\alpha_W,\theta;z_W):=\sum_{j\in\mathcal J_W}q_{1,j}(\alpha_j,\theta;z_j)\), with an empty sum interpreted as zero. The vectors \(z_S\) and \(z_W\) collect the corresponding fixed covariates.
For each weak study, let \(\mu_j^W\) be a finite retained-space dominating measure with full support on \(\mathcal W_j\), obtained by restricting or reparametrizing the original base measure when \(\mathcal W_j\) is full-dimensional and by using the chosen coordinate, Hausdorff, or other retained-space dominating measure when \(\mathcal W_j\) is lower-dimensional. Let \(p_j^W\) denote the retained-space baseline prior density induced by \(\pi_j\) on \(\mathcal W_j\) in the same restriction, reparametrization, or coordinate chart, and set \(\mu_W(d\alpha_W):=\prod_{j\in\mathcal J_W}\mu_j^W(d\alpha_j)\) and \(p_W^0(\alpha_W):=\prod_{j\in\mathcal J_W}p_j^W(\alpha_j)\). Thus \(p_W^0\mu_W\) is the unpooled baseline prior measure on the retained weak coordinates. The substantive hierarchy enters through \[ r_\lambda(\alpha_W) = \int_\Theta \exp\!\left( \lambda\left[q_{1,W}(\alpha_W,\theta;z_W)+q_{1,S}(\alpha_{S0},\theta;z_S)\right] \right) \pi_\theta(d\theta). \] The corresponding hierarchy-induced prior path on weak values is \[ \pi_\lambda(d\alpha_W) = p^\alpha_\lambda(\alpha_W)\mu_W(d\alpha_W), \qquad p^\alpha_\lambda(\alpha_W) = \frac{p_W^0(\alpha_W)r_\lambda(\alpha_W)} {\int_{\mathcal W}p_W^0(u)r_\lambda(u)\mu_W(du)} . \] The notation separates the prior measure from its density: \(\pi_\lambda\) is the probability measure on weak values, while \(p^\alpha_\lambda\) is its density in \(\alpha_W\)-coordinates. They are not different priors. Both are induced by the lower-level baseline prior, the upper-level prior \(\pi_\theta\), the pooling objective \(q_1\), and the pooling strength \(\lambda\). In local coordinates \(h=\alpha_W-\alpha_W^*\), we write \(p_\lambda(h)=p^\alpha_\lambda(\alpha_W^*+h)\) with respect to the translated retained-space measure. When \(q_1\) rewards similarity across studies, increasing \(\lambda\) shifts prior mass toward weak values that better satisfy the pooling relation with the other weak studies and the strong-study limiting values. In empirical terms, precisely estimated sites can affect the limiting weights placed on weakly estimated site-level treatment effects, IV slopes, production elasticities, or threshold values, without changing the weak-study moment criterion.
The finite-sample version replaces the limiting strong-study values by the unpooled strong-study estimates. Let \(\widehat\alpha^{\mathrm{un}}_{S,n}:=(\widehat\alpha_j^{\mathrm{un}})_{j\in\mathcal J_S}\), and set \[
\]
Under this definition, we have the following.
Proposition (ref) says that the hierarchy is a well-behaved way to tilt the weak experiment. The hierarchy need not be literally correct as a model for how the studies were generated. What matters is that, after the weak coordinates are retained, each value of \(\lambda\) gives a proper prior on \(\mathcal W\), and small changes in \(\lambda\) lead only to small changes in prior averages. The strong studies enter this prior through their first-order limits, so replacing those limits by the feasible strong-study estimates does not change the induced prior path asymptotically, uniformly over \(\lambda\in\Lambda\). The finite-sample reduced quasi-posterior uses \(Q_n\), while the limiting quasi-posterior uses \(Q\): \[
\] Let \(\Sigma(\alpha,\tilde\alpha):=\operatorname*{Cov}(\mathbb G(\alpha),\mathbb G(\tilde\alpha))\). When \(W(\alpha_W)=\Sigma(\alpha_W,\alpha_W)^{-1}\), the limiting quasi-posterior is the continuously updated weak-GMM quasi-posterior. Appendix (ref) gives the corresponding Andrews--Mikusheva diffuse-prior result: under the uniform diffuse-likelihood condition stated there, after absorbing the data-independent determinant factor into the structural prior, \(\Pi_\lambda(\cdot\mid g)\) is the diffuse-nuisance-prior limit of proper Andrews--Mikusheva Bayes posteriors. Using \(\pi_\lambda\), rather than the plug-in path, is harmless at this order because Proposition (ref) gives uniform total-variation convergence. Let \(T^{W}_{n,\lambda}:=\int_{\mathcal W}\alpha_W\,\Pi_{n,\lambda}(d\alpha_W\mid X,z)\) and \(t_\lambda(g):=\int_{\mathcal W}\alpha_W\,\Pi_\lambda(d\alpha_W\mid g)\). Within the reduced weak experiment, the unpooled quasi-posterior mean is \(T^{W}_{n,0}\) when \(\pi_0\) is proper; the usual unpooled weak-GMM point estimator is any approximate minimizer of \(Q_n\).
Our main result for weak studies is that the QBHM quasi-posterior is asymptotically equivalent to the limit experiment, when we look at the marginal distribution for weakly identified studies.
This proposition links the original QBHM to the reduced weak-GMM experiment. The convergence is strong enough to transfer the posterior-mean rule from the reduced weak experiment back to the feasible QBHM weak marginal. Let \(d_W:=\sum_{j\in\mathcal J_W}d_j\), the dimension of \(\alpha_W\). This transfer is shown by the following.
The convergence in Theorem (ref) is uniform over \(\lambda\in\Lambda\). Hence the same weak limit applies not only at any fixed choice of \(\lambda\), but also along data-dependent choices \(\hat\lambda_n\in\Lambda\). This is useful for the reporting strategy in Section (ref): one can display posterior means over a range of hierarchy strengths, or select a value of \(\lambda\) using the data, without changing the first-order weak asymptotic approximation.
The limit rule \(t_\lambda(g)\) has a direct Bayes-risk interpretation in the weak experiment. Once the weak-GMM process \(g\) has been observed, \(\Pi_\lambda(\cdot\mid g)\) is the quasi-posterior distribution over weak values, and its mean is the posterior squared-loss rule. For a reported linear estimand, let \(B\) be a fixed matrix; for example, \(B\) may select one component, a subset of components, a contrast across studies, or the full vector. Let a decision rule be a measurable map \(t:g\mapsto t(g)\in\mathbb R^{d_W}\), and use loss \(L_B(a,\alpha_W)=\|B(a-\alpha_W)\|^2\). If \(\mathsf M\) is any probability law on weak-process sample paths for which the quasi-posterior is well defined, define the integrated posterior risk \[ IPR_{\lambda,B,\mathsf M}(t) = \int \left[ \int_{\mathcal W}\|B(t(g)-\alpha_W)\|^2\Pi_\lambda(d\alpha_W\mid g) \right] \mathsf M(dg), \] interpreted as an extended nonnegative integral, so the value may be \(+\infty\).
Proposition (ref) gives the posterior squared-loss meaning of the pooled weak-GMM quasi-posterior. Combined with Proposition (ref) and Theorem (ref), it gives a decision-theoretic justification for the use of the QBHM estimator: for each fixed \(\lambda\), the weak marginal of the feasible QBHM estimator is asymptotically equivalent to a Bayes rule under the hierarchy-induced weak-limit prior \(\pi_\lambda\) and squared loss. In the continuously updated case, Proposition (ref) further shows that, under the diffuse-likelihood conditions stated there, this weak-limit posterior can be obtained as an Andrews--Mikusheva diffuse-nuisance-prior limit. Corollary (ref) in Appendix (ref) generalizes the result to the case where the same study can contain both strongly and weakly identified components. The conclusion is unchanged: fixed pooling does not alter the first-order law of genuinely regular coordinates, while weak coordinates retain the nondegenerate shrinkage comparison.
The previous section gives the posterior decision-theoretic interpretation of the QBHM estimator in the weak-GMM limit experiment. This section asks the pointwise frequentist question: for a fixed weak-GMM state \((\alpha_W^*,m)\), when does pooling lower MSE relative to an unpooled estimator? An applied version of this is essentially asking: when does borrowing from related sites reduce repeated-sampling error for a weakly estimated treatment effect. As indicated above, the answer comes from a bias-variance trade-off: pooling helps when the variance reduction from shrinkage exceeds the bias introduced by the pooling relation. The aim of this section is to formalize this trade-off for general QBHMs, and provide some more interpretable conditions in salient special cases. Throughout, \(\alpha_W\) denotes the weak-study vector from Section (ref), and \(\mathcal W\) denotes its compact retained weak space. The state restriction \(m(\alpha_W^*)=0\) means that the candidate weak value \(\alpha_W^*\) is compatible with the first-order drift of the weak moments. Appendix (ref) gives the corresponding formulas for profiled weak coordinates when a single study contains both regular and weak components.
Proposition (ref) is a posterior risk statement: after observing the weak-GMM process \(g\), the quasi-posterior mean is optimal for posterior squared loss under the hierarchy-induced weak-limit prior. Here, by contrast, we consider repeated-sampling MSE at a fixed weak-GMM state \(e=(\alpha_W^*,m)\), with \(m(\alpha_W^*)=0\). Thus \(\mathbb E_m\) denotes expectation over draws of the weak-GMM limit process \(g\), holding fixed both the true weak value \(\alpha_W^*\) and the drift \(m\). The distinction is exactly analogous to the distinction between Proposition (ref) and Corollary (ref) in the normal-normal example: the former averages risk using the posterior distribution, whereas the latter evaluates frequentist risk at a fixed underlying state. For any measurable rule \(t:g\mapsto t(g)\in\mathbb R^{d_W}\), and for any fixed matrix \(B\), define \[ M_B(t;\alpha_W^*,m)=\mathbb E_m\|B(t(g)-\alpha_W^*)\|^2 . \] This is the pointwise asymptotic MSE of the reported estimand \(B\alpha_W\). The choice of \(B\) records the estimand being reported: \(B=I_{d_W}\) gives the vector MSE, a row vector \(B=b'\) gives the scalar MSE for \(b'\alpha_W^*\), and a coordinate-selection matrix gives the MSE for selected components. The word pointwise is important because the risk does not average over the hierarchy, over possible values of \(\alpha_W^*\), or over a population distribution of studies.
The finite-sample analogue is \(\mathcal R_{n,\lambda}=\mathbb E_{P_{n,f}}\|B(T_{n,\lambda}-\alpha_W^*)\|^2\), where \(T_{n,\lambda}\) is the feasible estimator path, for example the weak-study component of the QBHM quasi-posterior mean. Proposition (ref) connects this finite-sample risk to the weak-limit risk. If \(B(T_{n,\lambda}-\alpha_W^*)\rightsquigarrow B(t_\lambda(g)-\alpha_W^*)\), uniformly over \(\lambda\in\Lambda\), and the squared errors are uniformly integrable, then \(\sup_{\lambda\in\Lambda}|\mathcal R_{n,\lambda}-M_B(t_\lambda;\alpha_W^*,m)|\to0\). Uniform integrability is a tail condition; it is automatic when the reported estimator path is bounded on a compact weak space and also follows from a common \(2+\delta\) moment bound. As such, a strict MSE improvement in the weak-limit experiment transfers to the original finite-sample sequence for all sufficiently large \(n\).
The comparison below is therefore a weak-study comparison. The comparator is the benchmark rule whose MSE is used as the reference point. In the QBHM comparison, the comparator \(t_0(g)\) is the unpooled weak-limit quasi-posterior mean obtained when \(\lambda=0\) and the baseline prior is proper, while \(t_\lambda(g)\) is the corresponding pooled weak-limit quasi-posterior mean. The same notation also allows \(t_0\) to denote another benchmark, such as an unpooled GMM rule or a leave-one-study-out rule. Define \(e_0(g)=B(t_0(g)-\alpha_W^*)\) as the benchmark estimation error for the reported estimand and \(d_\lambda(g)=t_\lambda(g)-t_0(g)\) as the change in the weak estimate caused by pooling.
Proposition (ref) says that pooling lowers MSE exactly when the covariance-type term \(2\mathbb E_m[e_0(g)'B d_\lambda(g)]\) is negative enough to offset the nonnegative cost term \(\mathbb E_m\|B d_\lambda(g)\|^2\). In words, pooling helps when the hierarchy tends to move the unpooled estimate back toward the truth, and it hurts when it tends to move the estimate away from the truth or when the pooling movement is too large relative to the error it corrects.
Two sufficient conditions are useful for interpreting the later results and are stated formally as Proposition (ref) in the appendix. First, it is trivially true that if the unpooled rule has infinite pointwise asymptotic MSE and the pooled rule has finite pointwise asymptotic MSE, then pooling strictly improves MSE. This is the extreme version of the variance-reduction logic, where the hierarchy regularizes an otherwise unstable weakly identified rule. Second, consider local coordinates \(h=\alpha_W-\alpha_W^*\). Suppose a sequence of hierarchy-induced priors concentrates around a local center \(h_c\), and let \(\delta_v(g)\) be the corresponding quasi-posterior mean. If \(\delta_v(g)\to h_c\) in \(L^2(P_m)\), then the MSE of the strongly pooled rule converges to \(\|Bh_c\|^2\). Therefore strong pooling improves on \(t_0\) whenever \[ \|Bh_c\|^2 < M_B(t_0;\alpha_W^*,m). \] The intuition is that very strong pooling makes the weak-study data almost irrelevant for the reported weak coordinate: the quasi-posterior mean collapses toward the target selected by the hierarchy. The remaining risk is then not a variance term from weak identification, but the squared distance between that target and the truth, measured after applying \(B\). Thus strong pooling is helpful only when replacing the noisy unpooled rule by the pooling target creates less squared error than the unpooled rule has on average. If the target is close to the true weak value, the variance reduction can dominate; if the target is badly centered, strong pooling simply replaces variance with bias. Lemma (ref) verifies the needed concentration step on compact local spaces, and the condition does not require the weak objective itself to become strongly curved.
The conditions above are general but abstract. The next two subsections make them more concrete: first by studying when a small amount of pooling helps relative to no pooling, and then by working through a scalar weak-IV example.
Starting from no pooling, the relevant local question is the sign of the first derivative of pointwise MSE with respect to \(\lambda\) at zero. The derivative separates the first movement of the weak quasi-posterior mean from the existing unpooled error, and in the quadratic case it becomes an explicit shrinkage-gain versus centering-cost comparison.
Work in local coordinates \(h=\alpha_W-\alpha_W^*\), and write \(\mathcal H=\mathcal W-\alpha_W^*\). In these coordinates the true weak value is the origin. Suppose \(\mathcal H\subset\mathbb R^{d_W}\) is compact with positive finite coordinate measure; lower-dimensional weak spaces are interpreted using a fixed coordinate parametrization. For a realized weak-GMM limit process \(g\), let \(Q_g(h)=Q(\alpha_W^*+h)\). Using the density notation introduced above, \(p_\lambda(h)\) is the local-coordinate density of the hierarchy-induced prior path \(\pi_\lambda\). The corresponding weak-limit quasi-posterior on local coordinates is the probability measure \(\nu_\lambda^g\) with density proportional to \(\exp[-Q_g(h)/2]p_\lambda(h)\) on \(\mathcal H\). Its mean \[ \delta_\lambda(g)=\int_{\mathcal H} h\nu_\lambda^g(dh) \] is the local error of the reported rule, so \(M_{\lambda,B}(\alpha_W^*,m)=\mathbb{E}_m\|B\delta_\lambda(g)\|^2\). At \(\lambda=0\), write \(\bar h(g)=\int h\nu_0^g(dh)\), \[ \Omega(g)=\int(h-\bar h(g))(h-\bar h(g))'\nu_0^g(dh), \] and, for \(P=P'\succeq0\), \[ \tau_P(g)=\int(h-\bar h(g))(h-\bar h(g))'P(h-\bar h(g))\nu_0^g(dh). \] Here \(\bar h(g)\) is the unpooled local mean error, \(\Omega(g)\) is the unpooled quasi-posterior variance, and \(\tau_P(g)\) is the third-central-moment correction that appears when the unpooled weak quasi-posterior is not centrally symmetric around \(\bar h(g)\).
The calculation only differentiates the hierarchy-induced prior path. It does not require differentiability or strong curvature of the weak objective \(Q_g\).
The score \(\ell_0\) is the first-order direction in which the hierarchy changes the local prior. Since \(p_\lambda\) is normalized and \(\ell_\lambda\) is uniformly bounded, dominated convergence gives \(\int_{\mathcal H} \ell_0(h)p_0(h)\,dh=0\). Thus the local prior score is automatically centered. If \(p_\lambda(h)\propto p_0(h)\exp[\lambda \tilde\ell_0(h)]\), with \(p_0\) bounded away from zero and infinity and \(\tilde\ell_0\) bounded, the normalized prior has score \[ \ell_0(h)=\tilde\ell_0(h)-\int_{\mathcal H} \tilde\ell_0(u)p_0(u)\,du. \] The centering constant has no effect on the covariance formulas below.
Under the compact local-coordinate conditions and Assumption (ref), Proposition (ref) in the appendix shows that, as \(\lambda\downarrow0\), \[ M_{\lambda,B}(\alpha_W^*,m) = M_{0,B}(\alpha_W^*,m) + 2\lambda \mathbb{E}_m\left[ (B\bar h(g))'B\operatorname{Cov}_{\nu_0^g}(h,\ell_0(h)) \right] + o(\lambda). \] The covariance term is the instantaneous movement of the quasi-posterior mean caused by the hierarchy. Small positive pooling lowers pointwise asymptotic MSE whenever this average directional derivative is strictly negative.
We can make this more explicit when the hierarchy induces a quadratic tilt of the local prior. Suppose the pooling contribution in local coordinates is \(q_1^W(h)=-(h-h_c)'P(h-h_c)/2\), where \(h_c\) is the local pooling center and \(P=P'\succeq0\). Since \(p_\lambda(h)\propto p_0(h)\exp[\lambda q_1^W(h)]\), the tilted density satisfies \(p_\lambda(h)\propto p_0(h)\exp[-\lambda(h-h_c)'P(h-h_c)/2]\). If \(p_0\) is the density of \(N(m_0,V_0)\), the exponent is \(-h'(V_0^{-1}+\lambda P)h/2+h'(V_0^{-1}m_0+\lambda P h_c)\), up to a constant, so \(p_\lambda\) is again normal, with precision \(V_0^{-1}+\lambda P\). Thus a normal prior and squared-loss pooling generate the quadratic path.
By the pooling manifold we mean the values that exactly satisfy the pooling relation: they are the values toward which the hierarchy shrinks and that receive no quadratic penalty. If \(P\) is nonsingular, the pooling manifold is the single point \(h_c\); if \(P\) is singular, it is an affine set. In the present local coordinates, the pooling manifold is \(\{h\in\mathcal H:P(h-h_c)=0\}\). In the fixed-center normal-normal hierarchy with \(P=I\), the manifold is the single point \(h_c\). In a common-mean normal hierarchy with \(J\) weak values and a diffuse common mean, the induced penalty is on deviations from the group average, \(P=I_J-J^{-1}\mathbf 1\mathbf 1'\), so the manifold is \(\{h:h_1=\cdots=h_J\}\); the hierarchy shrinks cross-study differences but leaves the common level unpenalized. For the quadratic path, \[ p_\lambda(h) = \frac{ \exp[-\lambda (h-h_c)'P(h-h_c)/2]p_0(h)} {\int_{\mathcal H}\exp[-\lambda (u-h_c)'P(u-h_c)/2]p_0(u)\,du}. \] The local prior score at no pooling is \[ \ell_0(h) = \left.\partial_\lambda\log p_\lambda(h)\right|_{\lambda=0} = c+h_c'Ph-\frac12 h'Ph, \] where the scalar \(c\) is the derivative of the normalizing constant. Differentiating only the prior weight gives, for each realized weak-GMM process \(g\), \[ \dot\delta_0(g) := \left.\partial_\lambda\delta_\lambda(g)\right|_{\lambda=0+} = \operatorname{Cov}_{\nu_0^g}(h,\ell_0(h)) = \Omega(g)P[h_c-\bar h(g)]-\frac12\tau_P(g). \] Define \[
\] Then \[ M_{\lambda,B}(\alpha_W^*,m) = M_{0,B}(\alpha_W^*,m) + 2\lambda[h_c'L_B(P)-A_B(P)] + o(\lambda). \] The term \(A_B(P)\) is the first-order shrinkage gain that would remain if the hierarchy were centered at the true local value \(h=0\). It is large when the unpooled rule has error in directions where the quasi-posterior still has dispersion to shrink. The term \(h_c'L_B(P)\) is the first-order centering cost from pulling toward \(h_c\) rather than toward the truth. Hence small positive pooling lowers asymptotic MSE whenever \(h_c'L_B(P)<A_B(P)\). If the unpooled weak quasi-posterior is centrally symmetric around \(\bar h(g)\), then \(\tau_P(g)=0\), and the gain margin reduces to \(\mathbb E_m[(B\bar h(g))'B\Omega(g)P\bar h(g)]\).
The same derivative calculation yields a uniform local improvement result, stated in \(\alpha_W\)-coordinates. For a weak-GMM limit experiment \(e=(\alpha_W^*,m)\), let \[ M_{\lambda,B}(e)=\mathbb{E}_m\|B(t_\lambda(g)-\alpha_W^*)\|^2 . \] Under the unpooled quasi-posterior \(\Pi_0(d\alpha_W\mid g)\), define the unpooled error and dispersion objects \[
\] Thus \(r_e(g)\) is the unpooled weak-limit error, \(\Omega_e(g)\) is the unpooled quasi-posterior variance, and \(\tau_{e,P}(g)\) is the same skewness correction in the original coordinates. Set \[
\] The scalar \(A_e(P)\) is the gain margin from shrinking in the penalized directions. The vector \(L_e(P)\) is the coefficient multiplying the centering error \(P(\alpha_W^*-\bar\alpha_W)\) in the expansion below. Finally, let \(D_{\mathcal W}=\sup_{u,v\in\mathcal W}\|u-v\|\) and \(K=\max(1,D_{\mathcal W}^3\|B\|_{\mathrm{op}}^2)\); this \(K\) is only a compactness bound used to state a uniform neighborhood around the pooling manifold.
A scalar fixed-center specialization makes the two terms in Theorem (ref) explicit. Let the weak parameter be scalar, take \(B=1\), and use a quadratic prior path centered at \(c\), so that \(P=1\) and \(\bar\alpha_W=c\). Write \(r(g)=t_0(g)-\alpha_W^*\), \(\omega(g)=\operatorname{Var}_{\Pi_0(\cdot\mid g)}(\alpha_W)\), and \(\kappa_3(g)=\int(\alpha_W-t_0(g))^3\Pi_0(d\alpha_W\mid g)\). The expansion becomes \[ M_\lambda-M_0 = -2\lambda \left[ \mathbb E_m\left[r(g)^2\omega(g)+\frac12r(g)\kappa_3(g)\right] +(\alpha_W^*-c)\mathbb E_m[\omega(g)r(g)] \right] +O(\lambda^2). \] If the unpooled weak quasi-posterior is locally symmetric, \(\kappa_3(g)=0\). If the hierarchy is correctly centered, \(c=\alpha_W^*\), small positive pooling lowers MSE whenever \(\mathbb E_m[r(g)^2\omega(g)]>0\): the rule has nonzero weak-limit error and the unpooled quasi-posterior has dispersion to shrink. If \(c\ne\alpha_W^*\), the extra term \((\alpha_W^*-c)\mathbb E_m[\omega(g)r(g)]\) is the centering-bias cost.
The two inequalities in Theorem (ref) have a direct interpretation. The first says that there must be residual dispersion for pooling to reduce: the unpooled weak quasi-posterior must still be spread out in directions that matter for the reported target \(B\alpha_W\), and the quadratic pooling term \(P\) must shrink those directions. This is the possible variance gain from pooling. The second says that the true weak value must not be too far from the values favored by the pooling relation. More precisely, it only has to be close in the directions that \(P\) actually penalizes. If \(P\) is singular, some directions are left alone. The set of values left unpenalized is the pooling manifold. In common-mean pooling, for example, the hierarchy shrinks cross-study deviations but does not shrink the common level, so the common level may be far from zero without violating the condition. Corollary (ref) gives a projection-based sufficient condition for the variance-gain part of the theorem, and the appendix records scalar fixed-center, coordinatewise ridge, and common-mean examples. In applications, large movements of \(\tilde\alpha_{n,\lambda}^{\mathrm{full}}\) along a reported \(\lambda\)-path should be read as evidence that the conclusion depends materially on the chosen pooling relation.
The scalar weak-IV example isolates a case in which the full comparison with \(\lambda>0\) can be written in closed form. Consider one study and suppress the study index. Let \(Y_i\) be the outcome, \(X_i\) a scalar endogenous regressor, and \(Z_i\) a scalar instrument, after any desired residualization on controls. The familiar just-identified IV model is \[ Y_i=\alpha_W^* X_i+u_i, \qquad X_i=\pi_n Z_i+v_i, \qquad \pi_n=\gamma/\sqrt n, \] with \(\mathbb E[Z_i u_i]=0\) and \(\mathbb E[Z_i v_i]=0\). Here \(\gamma\) is the example-specific local first-stage coefficient. The sequence \(\pi_n=\gamma/\sqrt n\) is the weak-first-stage sequence: the first stage is local to zero, so the IV slope is not pinned down at the usual \(\sqrt n\) scale.
For a candidate slope \(a\), the just-identified IV/GMM moment is \(G_n(a)=n^{-1/2}\sum_{i=1}^n Z_i(Y_i-aX_i)\). Writing \(h=a-\alpha_W^*\), we have \(G_n(\alpha_W^*+h)=G_{u,n}-G_{x,n}h\), where \(G_{u,n}=n^{-1/2}\sum_{i=1}^n Z_i u_i\) and \(G_{x,n}=n^{-1/2}\sum_{i=1}^n Z_i X_i\). Under the weak-IV sequence, \((G_{u,n},G_{x,n})\rightsquigarrow(G_u,G_x)\). Thus the weak-limit moment is \(M(h)=G_u-G_xh\), with scalar GMM weight one. The unpooled IV weak-limit rule solves \(M(h)=0\), giving \(H=G_u/G_x\) when \(G_x\ne0\). An induced normal pooling prior \(h\sim N(h_c,\lambda^{-1})\), with scalar pooling precision \(\lambda>0\), gives the pooled quasi-posterior mean \(\Delta_\lambda=(G_xG_u+\lambda h_c)/(G_x^2+\lambda)\). Here \(\Delta_\lambda\) is the weak-limit error of the pooled slope estimator, so the corresponding slope is \(\alpha_W^*+\Delta_\lambda\). For the conditional comparison, let \(b(G_x)=\mathbb{E}_m(G_u\mid G_x)\) and \(\sigma^2(G_x)=\operatorname{Var}_m(G_u\mid G_x)\). Conditional on \(G_x\ne0\), \[ \mathbb{E}_m(H^2\mid G_x)=\frac{b(G_x)^2+\sigma^2(G_x)}{G_x^2}, \qquad \mathbb{E}_m(\Delta_\lambda^2\mid G_x) = \frac{G_x^2\sigma^2(G_x)+[G_xb(G_x)+\lambda h_c]^2}{(G_x^2+\lambda)^2}. \] Subtracting the pooled conditional MSE from the unpooled conditional MSE gives \[ \mathbb{E}_m(H^2\mid G_x)-\mathbb{E}_m(\Delta_\lambda^2\mid G_x) = \frac{\lambda D_\lambda(G_x)}{G_x^2(G_x^2+\lambda)^2}, \] where \[ D_\lambda(G_x) = (2G_x^2+\lambda)(b(G_x)^2+\sigma^2(G_x)) -2G_x^3b(G_x)h_c -\lambda G_x^2h_c^2 . \]
Uniform improvement over a set \(\Lambda_1\) of positive pooling precisions requires \(\inf_{\lambda\in\Lambda_1}D_\lambda(G_x)>0\). The conditional mean \(b(G_x)\) records remaining local drift in the structural reduced-form component after conditioning on the weak first stage. In the orthogonalized case, \(b(G_x)=0\), as in the canonical Gaussian weak-IV limit when the structural reduced-form component is uncorrelated with the first-stage limit, or after replacing it by the residual from its linear projection on \(G_x\). The sign condition then reduces to \[ h_c^2<\sigma^2(G_x)(G_x^{-2}+2\lambda^{-1}), \] with a positive margin uniformly over \(\lambda\in\Lambda_1\) for uniform improvement on \(\Lambda_1\). The left side is the squared centering error of the hierarchy. The right side is the weak-identification variance gain from replacing \(G_x^{-1}\) by the ridge-stabilized factor \(G_x/(G_x^2+\lambda)\). When \(G_x\) is close to zero, the allowable centering error is large because unpooled IV is very noisy; when \(G_x\) is large, the allowable centering error is smaller because the unpooled ratio is already stable.
The finite-MSE part of Proposition (ref) allows \(G_x^{-1}\) to have infinite second moment while requiring the ridge-stabilized numerator \(G_xG_u\) to remain square integrable. Positive pooling then regularizes the weak-IV ratio uniformly on compact subsets bounded away from \(\lambda=0\).
The MSE results have a direct implication: a QBHM should not be presented as a single pooled estimate indexed by an unexplained value of \(\lambda\). The numerical value of \(\lambda\) is interpretable only after the normalization of \(q_1\) has been fixed. In the normal--normal example, \(\lambda\) determines the effective weight \(c_{j,n}(\lambda)\) on the unpooled group summary in Proposition (ref). Equivalently, when \(q_1\) is a Gaussian least-squares term with scale \(\tau_0^2\), the hierarchical precision is \(\lambda/\tau_0^2\). Changing \(\lambda\) and changing the variance scale inside \(q_1\) therefore change the same practical object: the amount of pooling. The role of \(\lambda\) in the general QBHM is both to handle \(q_1\) functions that do not have a natural “variance” term and to emphasize that researchers should report results over a grid of \(\lambda\) values, rather than at a single chosen value.
In applications, this should be a grid \(\Lambda_G\) that gives a range of strong to weak pooling under the chosen normalization of \(q_1\). A dense grid near zero is useful when small amounts of pooling are likely to matter, as highlighted by Theorem (ref). Tables should report marginal quasi-posterior means and quasi-posterior intervals over the grid, together with the exact normalization of the pooling term and the prior distribution used. Ideally, results should be robust to changes in \(\lambda\) and the prior, since these are chosen by the researcher. Theorem (ref) suggests that this should be true for strongly identified studies, but for weakly identified studies, both the prior and \(\lambda\) remain asymptotically relevant. It may nevertheless be the case that this does not matter much in practice, and reporting these checks can make the results more convincing to others. If the study's parameter changes substantially with \(\lambda\) and/or the prior, a researcher should note this and try to provide an empirical basis for their main choices, such as genuine prior knowledge from earlier studies, empirical Bayes for the prior, or cross-validation for \(\lambda\).
In such settings, cross-validation can provide a heuristic way to choose a good value for \(\lambda\). It may also be useful if a researcher wants to present one main set of results for policy reasons, provided that they still report results over several different values of \(\lambda\). For leave-one-study-out cross-validation, fit the QBHM on the \(J-1\) studies other than \(j\), and use that fit to predict an object in the left-out study. Let \(x_{-j}\) and \(z_{-j}\) collect the data and covariates from all studies except \(j\), and let \(\Pi^{\theta}_{n,\lambda,-j}(d\theta\mid x_{-j},z_{-j})\) be the marginal quasi-posterior for \(\theta\) computed without study \(j\). Define \[ w_{j,\lambda,-j}(a_j) =\int \exp[\lambda q_{1,j}(a_j,\theta;z_j)] \Pi^{\theta}_{n,\lambda,-j}(d\theta\mid x_{-j},z_{-j}). \] The induced predictive law for the left-out group parameter is \[ \Pi^{\mathrm{pred}}_{\lambda,-j}(d a_j\mid x_{-j},z) = \frac{w_{j,\lambda,-j}(a_j)\pi_j(d a_j)} {\int w_{j,\lambda,-j}(\tilde a_j)\pi_j(d\tilde a_j)} . \] In order to perform cross-validation, we need a way to evaluate how well the predicted law performs. Fortunately, many structural applications already provide such a criterion. For example, it is common to evaluate a model using moments that were not used to estimate its parameters, or more generally to assess whether it matches features of the data that it was not explicitly designed to match. This idea can be formalized through a validation criterion \(\mathcal V_j(a_j;x_j,z_j)\), such as a collection of validation moments, where smaller values indicate better predictive performance on held-out data or validation moments. Then \[ \operatorname{CV}(\lambda) =\sum_{j=1}^J \big \| \, \mathbb{E}_{\Pi^{\mathrm{pred}}_{\lambda,-j}} [\mathcal V_j(a_j;x_j,z_j)] \, \big\|, \qquad \hat\lambda_{\mathrm{CV}}\in\operatorname*{arg\,min}_{\lambda\in\Lambda_G}\operatorname{CV}(\lambda). \]
The selected value \(\hat\lambda_{\mathrm{CV}}\) is a tuning summary for the reported grid. Since \(J\) is fixed, we should not treat \(\hat\lambda_{\mathrm{CV}}\) as consistently learning some notion of the “oracle \(\lambda\).” However, we would generally expect the value selected by cross-validation to be more informative when \(J\) is large and \(\mathcal V_j(a_j;x_j,z_j)\) is an accurate validation criterion for the model. That said, a researcher should report results over a range of \(\lambda\) for robustness purposes.
Statistical inference under strongly identified studies is handled in the usual way. Due to Theorem (ref), the QBHM estimator is asymptotically equivalent to the GMM estimator, and so is asymptotically normal. Theorems 3 and 4 of ChernozhukovHong2003 then apply verbatim; under a generalized information inequality\footnote{See section 4 of ChernozhukovHong2003 for more information about what this is and how it can be guaranteed to hold. This can be done by appropriately constructing the weighting matrix in a GMM problem.}, draws from the quasi-posterior itself can be used to estimate the asymptotic variance. However, testing under weak identification is more difficult, and requires us to consider the weak limit discussed in Section (ref). If a researcher is concerned that some weakly identified parameter may be part of their hypothesis, they should test using the procedure below, rather than relying on draws from the quasi-posterior.
In that limit, testing a candidate weak value \(a_0\in\mathcal W\) corresponds to \(H_{0,a_0}:m(a_0)=0\): when \(a_0\) is the unique zero of \(m\), this is the usual point null \(\alpha_W^*=a_0\), while more generally it is the moment-validity null that is inverted to form a weak-identification-robust confidence set. In this subsection, \(\Sigma\) denotes the covariance function of the Gaussian weak moment process, not the GMM weight \(W\) used in the objective. Let \(\Sigma_{00}=\Sigma(a_0,a_0)\) be nonsingular, set \(V_{a_0}(\alpha_W)=\Sigma(\alpha_W,a_0)\Sigma_{00}^{-1}\), and define the residual path \(h_{a_0}(\alpha_W)=g(\alpha_W)-V_{a_0}(\alpha_W)g(a_0)\), the component of the process left after the Gaussian projection on \(g(a_0)\). Under \(H_{0,a_0}\) in the Gaussian weak-GMM limit experiment, \(g(a_0)\sim N(0,\Sigma_{00})\), and the Gaussian projection construction makes \(g(a_0)\) independent of the residual path \(h_{a_0}(\cdot)\). Hence, conditional on the observed residual path \(h_{a_0}\), the null law of the full process is generated by \(g^*(\alpha_W)=h_{a_0}(\alpha_W)+V_{a_0}(\alpha_W)\xi^*\), where \(\xi^*\sim N(0,\Sigma_{00})\). This conditional-inference argument follows AM2016conditional: the hierarchy can enter the test statistic or the alternative weighting, while the conditional null simulation supplies the critical value. For a statistic \(T_\lambda\), let the randomized conditional critical rule be
where \(c_{\zeta,\lambda}(h)\) and \(\rho_{\zeta,\lambda}(h)\in[0,1]\) are chosen so that the conditional null rejection probability is \(\zeta\). If the conditional null distribution is continuous at the critical value, the randomization term is immaterial.
The test above is one which would be valid in the weak limit, not one which is directly implementable. In practice, we would use a feasible version of this test. This uses a covariance-kernel estimator \(\widehat\Sigma_n\) for this inference step, replaces \(g\) and \(\Sigma\) by \(g_n\) and \(\widehat\Sigma_n\), and forms \[ h_{n,a_0}(\alpha_W) = g_n(\alpha_W) - \widehat\Sigma_n(\alpha_W,a_0) \widehat\Sigma_n(a_0,a_0)^{-1} g_n(a_0). \] The conditional simulation uses \[ g_n^*(\alpha_W) = h_{n,a_0}(\alpha_W) + \widehat\Sigma_n(\alpha_W,a_0) \widehat\Sigma_n(a_0,a_0)^{-1}\xi_n^*, \qquad \xi_n^*\sim N(0,\widehat\Sigma_n(a_0,a_0)). \] Proposition (ref) in the appendix records sufficient covariance-consistency, continuity, and no-boundary conditions for this test to have valid asymptotic size. Confidence sets are then obtained by inverting the conditional tests over candidate values \(a_0\). If a researcher is interested in testing \(B\alpha_W\) for some matrix $B$, rather than $\alpha_W$ itself, a valid confidence set is obtained by projection: if \(C_{n,W}\) is the inverted confidence set for \(\alpha_W\), then \(C_{n,B}:=\{B\alpha_W:\alpha_W\in C_{n,W}\}\) has at least the coverage of \(C_{n,W}\).
The hierarchy can also be used to form weighted-average-power (WAP) objective functions, which average power over alternatives using weights induced by the pooling relation AM2022. A hierarchy-weighted conditional quasi-likelihood-ratio statistic, formed from the weak-GMM quasi-likelihood under the null and over alternatives, puts more weight on alternatives that are close to the pooling relation when those alternatives are substantively central. Proposition (ref) in the appendix gives the conditional WAP calculation, and the main reporting distinction is that \(\lambda\) controls point-estimation shrinkage and can weight alternatives for power, while weak-identification-robust confidence sets deliver coverage.
For hypothesis that contain both strongly and weakly identified components, Appendix (ref), especially Corollary (ref), gives the separation between the root-\(n\) regular rate and the weak component's \(O(1)\), nondegenerate limit. Standard inference for strongly identified components applies to the strong block, while weak-identification-robust confidence sets apply to the weak block. A conservative set for a linear estimand on the original, unscaled parameter level, \(R_S\beta^S+R_W\alpha_W\), can be formed by combining a standard strong confidence set \(C_{n,S}(1-\eta_S)\) with a weak-identification-robust confidence set \(C_{n,W}(1-\eta_W)\), where \(\eta_S+\eta_W=\eta\). Projecting \(\{R_S\beta+R_W\alpha_W:\beta\in C_{n,S}(1-\eta_S),\,\alpha_W\in C_{n,W}(1-\eta_W)\}\) then gives a set with asymptotic coverage at least \(1-\eta\) under the joint strong/weak convergence conditions.
This section studies the finite-sample behavior of the QBHM estimator in an IV production-function design with ten groups. The design varies two features: first-stage strength and within-group sample size. Group parameters are drawn from a common distribution, so the pooling relation is informative but not exact. We compare the QBHM with unpooled GMM, using the same group-specific objective function in both cases. The results show the expected pattern: when instruments are weak or groups are small, pooling reduces dispersion and lowers average mean squared error (MSE) across sites. When within-group information is strong, the unpooled and pooled estimators are close.
The model estimates a Cobb--Douglas production function with capital as the variable input. The instrument is an exogenous treatment-induced shock to capital, in the spirit of de2008returns and fafchamps2014microenterprise. To ground the simulation, the data-generating process uses micro data from eight randomized controlled trials: CaiSzeidl2024, EggerEtAl2022, FinkJackMasiye2020, crepon2015estimating, augsburg2015impacts, attanasio2015impacts, karlan2019debt, and bari2024asset. Each study contains a treatment that shifts firm capital, and from these data we use three variables: a study identifier, pre-treatment capital, and treatment status. For each synthetic group, we sample a source study with replacement, sample observations from that study with replacement, and add lognormal jitter to capital. This preserves realistic cross-sectional capital distributions while allowing us to choose the number of groups and within-group sample sizes. Post-treatment capital is generated as:
Changing \(\delta_T\) and \(\sigma_\nu^2\) changes the first-stage strength of treatment as an instrument for post-treatment capital, with smaller \(\delta_T\) and larger \(\sigma_\nu^2\) producing weaker instruments.
Each synthetic group \(j\) is assigned production parameters drawn independently from the hierarchical distributions
Pre- and post-treatment revenue are generated from \(Y_{jit}=a_{1j}K_{jit}^{\rho_j}\exp(\varepsilon_{jit})\), where \(\varepsilon_{jit}\) is drawn independently across pre- and post-treatment periods from a Student-\(t\) distribution. In the moment objective function below, \(Y_{ji}:=Y_{ji1}\) and \(K_{ji}:=K_{ji1}\) denote post-treatment revenue and capital.
We apply the following QBHM to the synthetic data. The within-group stage uses an affine log-linear IV/GMM parameterization of the kind analyzed in the weak-IV example. Write \(\ell_j:=\log a_{1j}\) and \(\alpha^{\mathrm{CD}}_j=(\ell_j,\rho_j)^\top\), where CD denotes Cobb--Douglas. The objective function is written so that larger values are better: \[ q_{2,j,n_j}(\alpha^{\mathrm{CD}}_j) = -\frac{n_j}{2}\, \hat g_j(\alpha^{\mathrm{CD}}_j)^\top \hat g_j(\alpha^{\mathrm{CD}}_j). \] Here \(\hat u_{ji}(\alpha^{\mathrm{CD}}_j)=\log Y_{ji}-\ell_j-\rho_j\log K_{ji}\), \(\hat g_j(\alpha^{\mathrm{CD}}_j)=n_j^{-1}\sum_{i=1}^{n_j}\zeta_{ji}\hat u_{ji}(\alpha^{\mathrm{CD}}_j)\), and \(\zeta_{ji}=(1,T_{ji})^\top\). The moment is affine in \(\alpha^{\mathrm{CD}}_j\), and reported summaries for \(a_{1j}\) use the transformation \(a_{1j}=\exp(\ell_j)\).
The Gaussian pooling term follows the hierarchical structure of the data-generating process, with pooling on $\ell_j$ for the affine coefficient and on $\operatorname{logit}(\rho_j)$ to keep the production elasticity inside $(0,1)$:
The QBHM quasi-posterior is sampled using Hamiltonian Monte Carlo as implemented in Stan, with capital centered and scaled within groups to improve numerical conditioning.
The comparison estimator is the unpooled GMM estimator obtained by maximizing the same group-specific objective function \(q_{2,j,n_j}(\alpha^{\mathrm{CD}}_j)\) group by group. The transformation \(a_{1j}=\exp(\ell_j)\) is used when reporting results on the original production-function scale.
We run simulations that vary instrument strength and group sample size. The simulation parameters and weakly informative hierarchical priors are reported in Appendix (ref), Table (ref). The priors are centered on reasonable but false values, and their standard deviations are broad relative to the data-generating process. To check sensitivity to the prior specification, we rerun the simulations using very diffuse priors centered at zero, with variances more than 150 times the true data-generating-process scale. The very diffuse priors change which value of \(\lambda\) performs best in this design, but do not change the main comparison with the unpooled estimator; the sensitivity results are reported in Appendix (ref).
Table (ref) reports the simulation results, with the top panel varying instrument strength. The “Very strong” scenario corresponds to \((\delta_T=0.5,\sigma_\nu=0.4)\) and an average realized first-stage \(F\)-statistic across groups and replications of \(3813.89\). The “Weak” scenario, \((\delta_T=0.06,\sigma_\nu=1.1)\), and the “Very weak” scenario, \((\delta_T=0.03,\sigma_\nu=1.4)\), have average \(F\)-statistics equal to \(8.34\) and \(2.06\), respectively.
The table reports average mean squared error for point estimates and simulation coverage of nominal 95% intervals, using central quasi-posterior intervals for the QBHM and Wald intervals for the unpooled estimator. The overall pattern is consistent with the theory: when instruments are strong, the unpooled and QBHM estimators are close because within-group information dominates the fixed pooling term. When instruments are weak, the QBHM estimator reduces dispersion enough to lower root mean squared error (RMSE) even after allowing for centering induced by the hierarchy. The sample-size experiments show the same mechanism, with the bottom panel using the “Very strong” instrument setting and varying the common group sample size \((n_j)\). When each group contains little information, pooling reduces dispersion and improves average MSE, while the gains from pooling are small when each group is precise. In addition, the results on coverage show the lessons of Section (ref): confidence intervals made by drawing from the quasi-posterior have good coverage when instrument strength is strong and group size is relatively large, but coverage gets generally worse when these do not hold. As such, for weak-identification-robust coverage in empirical work, Section (ref) recommends weak-identification-robust confidence sets.
The \(\lambda\)-grid gives the corresponding sensitivity check: in the “Very strong” scenario with large sample size, the unpooled estimator has lower RMSE at \(\lambda=1\), because strong within-group information leaves little variance for the hierarchy to remove. As \(\lambda\) decreases, the QBHM estimator moves toward the unpooled estimator. Full results for the grid are reported in Appendix (ref), Table (ref). The \(\lambda\)-grid result is the finite-sample simulation analogue of the theoretical bias--variance tradeoff: more pooling can lower dispersion when within-group information is weak, but it can increase MSE when the hierarchy adds centering bias to a strongly informative group objective function. In this design, the central QBHM quasi-posterior intervals tend to be conservative as finite-sample simulation summaries, while the unpooled estimator Wald summaries undercover.
We illustrate the QBHM approach using microenterprise production functions from randomized controlled trials in the capital-drop literature. The empirical exercise uses micro data from four studies: EggerEtAl2022, denoted GE; bari2024asset, denoted AB; de2008returns, denoted RC; and fafchamps2014microenterprise, denoted FY. All four studies analyze microenterprises, include a treatment that induces a capital shock, and record firm capital as a primary or secondary variable. The purpose is to show how the estimator behaves in a nonlinear grouped objective function with unstable study-specific features.
The model is based on the production-function component of the microenterprise poverty-trap model in banerjee2019can.\footnote{banerjee2019can is a working paper at the time of writing and the dataset is not yet publicly available.} Entrepreneurs choose between two technologies. One has diminishing returns to capital, while the other has constant returns but requires a fixed cost to be paid. This choice is represented by a production function with a Cobb--Douglas component at low levels of capital and a linear component at high levels of capital. A nonconvex production function is necessary for a poverty trap in this class of models, and the location of the nonconvexity determines both whether a trap exists and how important it is.
Estimating the production function is challenging because the point of nonconvexity is unknown, depends on the parameters of the function, and may lie outside the support of observed capital for some groups. When the switching threshold lies on the edge of, or outside, the observed support for a group, the model parameters are weakly identified or not identified. banerjee2019can addresses this difficulty in a well-powered experimental setting by assuming the existence of a trap and selecting from a discrete grid of parameter values using a GMM objective function. Because the QBHM approach pools information across studies while preserving study-level heterogeneity, the application makes the theory's shrinkage mechanism visible in a nonlinear moment problem.
The nonconvex production function we estimate is the following:
Here \(Y_{jit}\) is observed revenue, \(K_{jit}\) is observed capital, and \(K_{\mathrm{thresh},j}\) is an unobserved fixed cost required to access the higher-return technology. The QBHM and unpooled estimators estimate the parameter vector \(\alpha^{\mathrm{NC}}_j=(a_{1j},a_{2j},\rho_j,K_{\mathrm{thresh},j})^\top\), where NC denotes nonconvex.
The estimation process is the same as in Section (ref), except that the residual \(\hat u_{ji}(\alpha^{\mathrm{NC}}_j)\) is computed from the nonconvex production function and the QBHM additionally assumes
The compact-support theory in Section (ref) treats parameter spaces as bounded, so these Gaussian specifications can be read as their restrictions to the finite parameter sets used for estimation; the display reports the underlying centers and scales.
To evaluate finite-sample performance in a setting close to the application, we run Monte Carlo simulations with a data-generating process whose structure and parameters mimic the applied setting. The details and results of the simulations are reported in Appendix (ref). In this application-calibrated simulation design, the QBHM has lower RMSE and better finite-sample simulation interval coverage than the unpooled estimator for all four parameters.
Both estimators recover the proportion of firms below the switching threshold, with RMSE for the trapped proportion ranging from 0.014 to 0.056 and a slight edge to the QBHM.\footnote{The proportion of trapped firms depends on the location of the nonconvexity, which in turn depends on multiple parameter values. It is possible for two models to have very different parameter estimates while still choosing the same nonconvexity location.} The result reflects the fact that most simulated groups have a switching threshold inside the observed support, so the threshold is disciplined by observed variation rather than by extrapolation.
We use $\lambda=1$ as the reported pooled reference because the pooling terms are normalized as Gaussian hierarchical log densities. We also report $\lambda=0$ as the unpooled reference and rerun the analysis with $\lambda=0.5$ and $\lambda=2$ to assess sensitivity. The sensitivity checks, reported in Appendix (ref), show that pooling affects estimates relative to $\lambda=0$, while estimates are stable across the positive pooling values $\lambda\in\{0.5,1,2\}$. The $\lambda$ values index alternative induced prior paths in the sense of Section (ref).
The empirical application does not select \(\lambda\) by cross-validation. Instead, it uses the prespecified grid \(\{0,0.5,1,2\}\) to make the normalization and sensitivity transparent. Stable conclusions across the grid are driven mainly by within-study moment information, while conclusions that appear only for larger $\lambda$ rely more heavily on the pooling relation. When cross-validation is implemented, the CV-selected tuning value should be chosen from the same grid using a prespecified left-out-study score, such as predictive GMM fit or accuracy for the switching threshold.
Both the QBHM and unpooled estimators are used to estimate the production-function parameters in the four study data sets. Parameter estimates are reported in Table (ref) in Appendix (ref). Using the production-function point estimates, we calculate the implied switching threshold for capital, where \(a_{1j} K^{\rho_j}=a_{2j}\!\left(K - K_{\mathrm{thresh},j}\right)\), and the proportion of firms in the sample operating below this threshold. Table (ref) reports the switching capital threshold and the proportion of firms below it for each of the four studies.
The estimator behavior is consistent with the weak-identification motivation, and we interpret the estimates as evidence of a poverty trap in the observed support only when the fitted nonconvexity switch lies within the study sample. If all observed firms are fitted with the Cobb--Douglas component, we treat the exercise as finding no within-support switch, rather than as precisely estimating a trap outside the support.
The QBHM estimates a within-support nonconvexity switch, and hence a poverty-trap pattern under the maintained production-function model, in one study, fafchamps2014microenterprise (FY). Approximately 80% of the sample lies below the switching threshold, and the estimated fixed cost creating the trap is \$446 in purchasing-power-parity terms. In the other three studies, the QBHM places the switching threshold above or near the edge of the study support. Care is needed when interpreting parameter values that fall outside the study support, as illustrated by the very high GE and AB switching-capital estimates. We interpret these results as evidence of no estimated switch within the support, not as precise estimates of distant economically meaningful thresholds. A switching threshold far outside the sample is consistent with the model behavior expected when no within-support switch is present.\footnote{In both studies the QBHM has little direct evidence about the linear-branch parameters of the production function, $(a_2,K_{\mathrm{thresh}})$, so those components are disciplined mainly by the pooling relation and the other studies.}
The comparison with unpooled estimation shows where pooling matters. The unpooled estimator agrees with the QBHM in the AB and RC studies in detecting no within-support switching threshold, although it assigns the observed support to a different branch of the production function. The unpooled estimator also estimates a within-support trap in FY, but at a lower capital threshold than the QBHM. It also estimates a within-support trap in EggerEtAl2022 (GE), while the QBHM places the switch outside or near the edge of the observed support. Taken as an implementation exercise, the estimates indicate a trap in FY, no trap in RC or AB, and a less stable conclusion in GE.
Although the model implies that treatment effects should be larger for firms raised above the threshold by the RCT intervention, the present empirical section does not estimate that heterogeneity; its role is to show how QBHM organizes an unstable nonlinear grouped objective function.
This paper develops the QBHM as a way to borrow information across related econometric studies without forcing all studies into the same model. Each study keeps its own objective function, instruments, controls, weights, fixed effects, and parameter dimension, with the hierarchy only added for components that the researcher believes are meaningfully comparable. With a fixed number of studies, the goal is not to learn one common cross-study parameter; instead, it is to learn the study-specific estimands while using the hierarchy as a transparent pooling relation.
The results imply that, when a study is already precise, fixed pooling has little first-order effect. When a study is noisy or weakly informative, pooling can materially affect estimates: it can reduce error if the hierarchy is well centered, but it can add bias if the pooling relation is wrong. The decision-theoretic result should be read in this operational sense: for weakly identified blocks, the feasible QBHM quasi-posterior mean converges to the weak-GMM Bayes rule under squared loss for the reported estimand. As such, there is a decision-theoretic foundation for recommending the use of the QBHM estimator, even when some studies may be weakly identified. Since $\lambda$ and prior choice may potentially affect results, empirical work should report estimates over a path of pooling strengths, not just one pooled estimate. A cross-validation choice of \(\lambda\), when used, should be treated as a reference point on that path. Findings that are stable across the path are mainly driven by within-study information; findings that appear only under strong pooling rely more heavily on the hierarchy. The simulations and application show the intended use: QBHM behaves much like unpooled estimation when the group evidence is strong, and is most useful when some group objectives are noisy, nonlinear, or weakly informative. Taken together, QBHM gives researchers a disciplined way to adaptively combine information from various studies, without making any strong distributional assumptions. Further applying this approach to substantive questions in economics remains an exciting direction for future research.