Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
110,936 characters · 15 sections · 108 citation commands
Plausible GMM: A Quasi-Bayesian Approach
Moment restrictions are commonly used in the identification and estimation of structural or causal parameters in empirical economics. Prominent examples include instrument exclusion conditions, unconfoundedness assumptions, parallel trend assumptions, and nonlinear moment restrictions imposed in structural models. Economists typically use institutional knowledge and economic reasoning to argue for the validity of these restrictions in settings with observational data. Based on these arguments, classical estimation and inference, such as estimation and inference based on generalized method of moments (GMM), then proceed under the maintained assumption that the posited moment restrictions hold exactly.
While the arguments employed to justify moment restrictions provide a basis for believing that the moment conditions are plausible, they are also generally debatable. That is, it is hard to know whether there are unmodeled sources of confounding or sources of misspecification that would result in moment conditions failing to hold exactly in any given empirical setting. Unfortunately, it is well-known that estimation and inference results obtained under the assumption that moment restrictions hold exactly can be substantially distorted in the sense of returning biased estimates and delivering unreliable conclusions. See, for example, hansen2001acknowledging, hansen2008robustness, hansen2010fragile, hall2003large, cheng2015uniform, CHEN2024105653 and hansen2021inference.
In this paper, we consider one approach to estimation and inference within a semiparametric structural model characterized by a set of moment restrictions, allowing for the possibility that the specified moment conditions do not hold exactly. We consider a setting where a researcher has access to an observable independent and identically distributed (i.i.d.) data stream $\{Z_t\}_{t=1}^{T}$ realized from unknown distribution $\mathbb{P}_{\mu_*}$.\footnote{While we maintain the assumption that the $Z_t$ are independent and identically distributed, some of our theoretical results hold more generally. We briefly comment on extensions to non-i.i.d. settings in Section (ref).} The researcher specifies a structural model, defined in terms of an economically meaningful $k$-dimensional parameter vector $\theta_*$, that restricts the joint distribution via moment conditions $m(\theta_*)=\mathbb{E}_{\mathbb{P}_{\mu_*}}[g(Z_{t},\theta_*)] = \mu_*$ for some $q \geq k$ dimensional vector $\mu_*$. Of course, informative inference about $\theta_*$ is impossible without beliefs about $\mu_*$. Classical estimation and inference proceed under the dogmatic prior belief that $\mu_* \equiv 0$.
Rather than adopt dogmatic prior beliefs, we conceptualize the notion that moment restrictions are plausible---but not known to hold exactly---by assuming the researcher is able to place an informative prior distribution over $\mu_*$. The use of a proper prior over $\mu_*$ allows informative inference about model parameters to proceed while relaxing the usual restriction that $\mu_* \equiv 0$. By concentrating this prior over 0, we capture the notion that a researcher subjectively believes the structural moment restrictions are likely to be correct. The spread and shape of the prior away from 0 further captures the researchers' beliefs about likely economically motivated possible deviations from the baseline structural model. Thus, the use of a proper prior distribution over $\mu_*$ provides a way for researchers to explicitly encode their subjective beliefs over the plausibility of their structural model.
Given that we wish to only leverage structural moment conditions and choose to conceptualize the plausibility of these moment conditions by using a proper prior distribution, it is natural to consider estimation and inference based on approximate or quasi-Bayesian posteriors (QBP), as in kim:lilgmm and chernozhukov2003mcmc.\footnote{In chernozhukov2003mcmc, estimators produced from quasi-Bayesian posteriors are referred to as Laplace-Type Estimators (LTEs). These approaches are also connected to “probably approximately correct inference,” e.g., catoni:learning.} QBPs provide a tractable approach to approximate Bayesian estimation and inference in traditional semiparametric moment condition models where $\mu_* \equiv 0$ is imposed; see, e.g., kim:lilgmm, gallant:inducedspace, fs:bayesmoment, and andrews2022optimal. Outside of the Bayesian motivation, chernozhukov2003mcmc demonstrate that these estimators have desirable frequentist properties within the moment condition framework when $\mu_* \equiv 0$. In addition, andrews2022optimal verifies that quasi-Bayes decision rules approximate Bayes' optimal decision rules within the weakly identified GMM framework.
In this paper, we extend the QBP framework to deal with settings where moment conditions are not assumed to hold exactly, i.e., to settings with non-dogmatic prior over $\mu_*$. We refer to estimation and inference within this setting as plausible GMM (PGMM). A central challenge in this setting is that $\theta_*$ and $\mu_*$ are not jointly identified, which implies that the impact of priors is not asymptotically negligible. We develop new technical results that address this challenge and show that, under suitable regularity conditions, key properties of the QBP framework extend to the PGMM setting. Building on andrews2022optimal, we verify that quasi-Bayes decision rules approximate Bayes optimal decision rules given the provided subjective priors. We also generalize the results of chernozhukov2003mcmc to verify that interval estimates from QBPs have a well-defined ex-ante notion of frequentist coverage under a sampling process where nature first draws a degree of misspecification from the subjective prior for $\mu_*$ and then the model realizes conditional on this draw as in conley2012plausibly and analogous to the coverage notion considered in andrews:pseudo. Finally, we provide novel large sample approximations for quasi-posterior distributions within this partially identified framework, allowing for the dimension of both $\theta_*$ and $\mu_*$ to increase with the sample size. These results can be viewed as new Bernstein-von Mises type theorems that explicitly account for additional terms that arise when dealing with misspecification.
We illustrate the use of QBPs with proper priors over the degree of misspecification, $\mu_*$, via two empirical examples. A cost of allowing for potential misspecification by considering non-dogmatic beliefs about $\mu_*$ is that inferential statements must be less precise than those obtained under dogmatic beliefs. The empirical applications demonstrate that one can still draw economically meaningful conclusions using our approach under what we consider to be sensible beliefs about model misspecification. The approach thus potentially enhances the credibility of the qualitative empirical results. We also use the empirical examples to discuss prior choice, illustrate the impact of prior choice on the resulting quasi-posteriors, and discuss sensitivity analysis focused on prior choice.
There is a large existing literature on sensitivity analysis and partial identification. Much of this research focuses on establishing formal frequentist guarantees for estimating and performing inference on the identified set. See, e.g., Canay_Shaikh_2017 and molinari:review for excellent reviews and norets2014semiparametric, kline2016bayesian, chib2018bayesian, liao2019bayesian, giacomini2021robust, giacomini2022robust, and kuangbayesian for approaches leveraging Bayesian methods.
Within this literature, our work is closely related to armstrong2021sensitivity. They consider moment condition models in which, under correct specification, $m(\theta_*) = 0$ for a true population parameter $\theta_*$. Misspecification is introduced by allowing $m(\theta_*) = C_T \neq 0$, where $C_T$ is unknown but constrained to lie in a known set; see also bonhomme2022minimizing. armstrong2021sensitivity focus on the case $C_T = c/\sqrt{T}$, where misspecification is local in the sense that it is of the same order as sampling uncertainty. Their framework, however, also accommodates settings in which misspecification does not vary with sample size. They develop a tractable frequentist approach to valid inference, show how sensitivity analysis can be conducted by varying the admissible set for $C_T$, and demonstrate that optimal GMM weighting matrices in this environment optimally trade off sampling variability against potential misspecification.
Our analysis delivers an analogous but conceptually distinct result from a quasi-Bayesian perspective. In particular, we show that QBPs are centered on a GMM estimator whose weighting matrix trades off moment precision with misspecification in a manner related to armstrong2021sensitivity. This trade-off emerges endogenously from the quasi-posterior when Gaussian priors with variance proportional to $1/T$ are imposed, rather than being imposed through a minimax or sensitivity criterion. While the resulting weighting structure differs in detail, this connection highlights how robustness considerations arise naturally in quasi-Bayesian inference under local misspecification. \textcolor{black}{We further show that, in the Gaussian limit experiment, the Bayes credible interval produced by our procedure under a two-point least favorable prior coincides with the robust confidence interval of armstrong2021sensitivity; see Supplementary Appendix Section (ref).} This equivalence provides a formal link between the two approaches and clarifies the sense in which the quasi-Bayesian method recovers existing robust frequentist procedures in the least favorable case. It also helps explain why, under alternative priors, our confidence intervals are typically narrower than those obtained from the procedure of armstrong2021sensitivity, as illustrated in the empirical application in Section (ref).
Another closely related paper is chen2018monte, who use simulation from quasi-posteriors to construct confidence sets with frequentist coverage guarantees for identified sets in general settings, including moment condition models. They illustrate their approach in moment inequality models by augmenting the model with an auxiliary parameter $\mu_*$, imposing support restrictions implied by the moment inequalities, and profiling out $\mu_*$ rather than placing a prior on it. While their framework could in principle be extended to other settings with support restrictions on $\mu_*$, including the environment considered here, our focus is on a fundamentally different inferential regime in which a prior is imposed on $\mu_*$. This distinction is not innocuous: the presence of a prior has important implications for posterior concentration and alters the asymptotic behavior of the quasi-posterior in partially identified models. Our resulting theory relies on arguments that differ from those in chen2018monte and yields insights that are not directly available under profiling-based approaches. These theoretical differences are illustrated through concrete examples in Supplementary Appendix Section (ref).
Our perspective is different from much of this previous work whose chief goal is establishing frequentist guarantees under partial identification as we wish to impose a proper subjective prior over $\mu_*$. That is, we mostly maintain a subjective Bayesian perspective as we believe there are settings where researchers will want to employ informative, subjective beliefs about potential misspecification. Our work thus aligns closely with the strand of Bayesian work on partial identification reviewed in gustafson:book. An element of this work is establishing posterior concentration results. We contribute to this literature by providing such concentration results within the semiparametric moment condition framework where the source of partial identification is uncertainty about the potential misspecification. These concentration results also allow us to consider frequentist properties of posterior summaries and thus complement the broader literature on partial identification and sensitivity analysis.
Within the Bayesian literature on partial identification and misspecification, our setup resembles chib2018bayesian. chib2018bayesian consider a Bayesian semiparametric moment condition model that includes an auxiliary parameter equivalent to $\mu_*$ to capture the misspecification of some moment conditions. However, chib2018bayesian focus on establishing posterior concentration on a well-defined pseudo-true value in the case of model misspecification, which requires that the number of free elements in $\mu_*$ is less than $q-k$. We instead allow for the possibility that all elements of $\mu_*$ are free, which precludes point identification of even a pseudo-true value and complicates establishing asymptotic concentration. Our formal results differ substantially in that priors have a non-negligible impact on asymptotic results, and posteriors do not generally concentrate on a unique pseudo-true value.
Our work also complements the recent contribution of andrews:pseudo, who study Bayesian decision-making in over-identified population minimum distance problems under potential misspecification. By considering appropriately constructed misspecification-invariant statistics, they provide an approach for obtaining interval estimates that have a notion of ex ante frequentist coverage under a class of distribution-invariant priors. Our approach applies in just-identified as well as over-identified settings under relatively general subjective priors over misspecification. Our motivation is chiefly Bayesian, but our approach provides a similar notion of ex ante frequentist coverage. We thus view the two approaches as complementary.
In summary, we develop a framework that treats potential misspecification in semiparametric moment condition models through a partial identification perspective induced by parameter over-parameterization. Rather than imposing the dogmatic benchmark that the moment restrictions hold exactly, we allow the researcher to place a proper prior on the degree of misspecification. This leads to a quasi-Bayesian approach to estimation, decision-making, and uncertainty quantification that remains informative under both global and local misspecification. We also show that informative prior restrictions can substantially sharpen the resulting identified set. More broadly, the quasi-posterior provides a tractable way to summarize uncertainty about both the parameter of interest and the extent of misspecification.
The remainder of the paper is organized as follows. In Section (ref), we more carefully discuss the main ideas and provide a convenient approximation result for the case when the prior for $\mu_*$ is taken to be Gaussian with precision proportional to the sample size---the case of “local misspecification.” We present the empirical illustrations in Section (ref), and we provide formal results in Section (ref). The Supplemental Appendix (indexed by the prefix “SA”) and the Online Materials (indexed by the prefix “OM”) provide additional details and proofs.
Suppose that we observe data $\{Z_{t}\}_{t=1}^{T}$ which are a realization from some unknown distribution $\mathbb{P}_{\mu_*}$. Suppose that we also have a posited structural economic model, which provides a set of moment restrictions on the distribution $\mathbb{P}_{\mu_*}$ indexed by a $q$-dimensional parameter $\mu_* \in \mathcal{M}$. Specifically, suppose the structural model implies a set of $q \geq k$ equations for a $k$-dimensional parameter $\theta \in \Theta$
such that there exists a vector $\mu_*$ corresponding to a target parameter $\theta_*$ satisfying
Of course, with no restrictions on the vector $\mu_*$, it is impossible to update beliefs about $\theta_*$ or the distribution $\mathbb{P}_{\mu_*}$ using the structural model. For any posited value of $\theta$ and distribution $\mathbb{P}_{\mu}$, we can always set $\mu = \mathbb{E}_{\mathbb{P}_{\mu}}[g(Z_{t},\theta)]$ such that the structural moment equation is satisfied.\footnote{Priors over $\theta_*$ and $\mathbb{P}_{\mu_*}$ produce restrictions over $\mu_*$, but the structural moment restriction adds no additional information if $\mu_*$ is left completely unrestricted.} Classical approaches to moment restriction models bypass this difficulty by assuming the vector $\mu_*$ is known to be a fixed, prespecified vector (without loss of generality $\mu_* \equiv 0$). This classical approach is equivalent to imposing the dogmatic prior that the researcher knows the structural moment equations hold exactly under $\mathbb{P}_{\mu_*}$---that is, the researcher has a dogmatic prior that the moment equations are correctly specified.
Unfortunately, it is hard to be fully confident that a set of structural moment restrictions hold exactly in many settings. For example, we may worry that there are unobserved confounding variables or that the functional form of the structural model is incorrect. We allow for departures from the dogmatic belief that the structural moment restrictions hold exactly by making use of a proper, non-degenerate prior distribution over $\mu_*$, denoted $\pi(\mu)$. The use of a proper prior over $\mu_*$ allows moment restrictions to be informative in updating beliefs about $\theta_*$ while falling short of imposing the often implausible restriction that moment restrictions hold exactly.
As a concrete example, consider the constant coefficient linear model
where $X_t$ is an observed variable with $\mathbb{E}_{\mathbb{P}_{\mu_*}}[X_t U_t] \neq 0$. Further, suppose we observe an additional variable $D_t$ that, based on economic reasoning or institutional knowledge, we believe satisfies the usual instrument exclusion restriction $\mathbb{E}_{\mathbb{P}_{\mu_*}}[D_t U_t] = 0$ for $t = 1,...,T$. Under this belief, we obtain the moment restriction $\mathbb{E}_{\mathbb{P}_{\mu_*}}[D_t U_t] = 0$ which can be used to identify the structural parameter $\theta_*$.
However, it is hard to know that the IV exclusion restriction holds exactly. For example, we might worry that there exists an unobserved confound, $M_t$, that covaries with both $Y_t$ and $D_t$ such that $U_t = M_t + V_t$, $\mathbb{E}_{\mathbb{P}_{\mu_*}}[D_t M_t] =\mu_* \neq 0$ and $\mathbb{E}_{\mathbb{P}_{\mu_*}}[D_t V_t] = 0$. Imposing the moment restriction $\mathbb{E}_{\mathbb{P}_{\mu_*}}[D_t U_t] = \mathbb{E}_{\mathbb{P}_{\mu_*}}[D_t (Y_t - \theta X_t)] = 0$ and solving for $\theta$ produces
Within the IV example, we might instead consider the restriction $\mathbb{E}_{\mathbb{P}_{\mu_*}}[D_t (Y_t - \theta_* X_t)] = \mu_*$ where we assume that $\mu_*$ is a fixed realization from a random variable $\mu$, e.g., $\mu \sim N(0,\sigma^2)$. Here, the assumed distribution captures the notion that the researcher believes the instrument is “close to” being valid in that the prior mass for $\mu$ is concentrated around 0. The distribution also encapsulates that the researcher believes it is incredibly unlikely that the moment restriction is perfect as $\{\mu = 0\}$ occurs with zero probability under such a distribution. Finally, the researcher can control beliefs about the strength of the unobserved confound via the prior variance, $\sigma^2$, while technically allowing for $\mu_*$ to be unbounded. That is, the proper prior over $\mu_*$ allows a well-defined and concrete description of the moment restriction being plausibly, but not certainly, satisfied.
To summarize, we are interested in a formalized version of a “plausible” moment restriction model characterized by parameters $(\theta,\mu)$ such that
and $\mu$ is governed by a prior distribution with density $\pi(\mu)$. We refer to $\mu$ as the “plausibility characteristic,” and we denote any root of the equation $m(\theta) = \mu$ as $\theta(\mu)$. For establishing some of the formal results in Section (ref), we will assume that $\pi(\mu)$ places strictly positive mass over a region $\Gamma$ such that solutions $\theta(\mu)$ exist for $\mu \in \Gamma$. This prior restriction is essentially trivially satisfied for any prior when $q = k$; see, e.g. hall2003large. However, satisfaction of this assumption is not guaranteed with $q > k$, suggesting that care should be taken in adding moment conditions about which a researcher has relatively weak prior beliefs unless the researcher is willing to use very diffuse priors.\footnote{From a frequentist perspective, armstrong2021sensitivity note that usual overidentification statistics can be used to infer lower bounds on the magnitude of $\mu$. andrews:pseudo consider an interesting different approach in overidentified settings that operationally inflates the size of confidence sets based on the magnitude of overidentification statistics.}
In the next section, we outline a quasi-Bayesian approach to perform inference on our main target parameter $\theta$.
We adopt a quasi-Bayesian approach to performing inference within the plausible moment restriction model. Let
be the average of $g(Z_t,\theta)$ against the empirical distribution at a given value of $\theta$. We can then define a continuous updating GMM-type criterion function for parameters $(\theta,\mu)$ as
for $\widehat{\Omega}_{T}(\theta)$ a positive definite matrix approximating $$\Omega(\theta)= \lim_{T\to \infty} \text{Var}\left(\sqrt{T}(\widehat{m}(\theta)-m(\theta))\right).$$ For example, it would make sense to use $$\widehat{\Omega}_T(\theta) = \frac{1}{T}\sum_{t=1}^{T} \left(g(Z_t, \theta)-\widehat{m}(\theta)\right)\left(g(Z_t,\theta)-\widehat{m}(\theta)\right)^{\top}$$ under the assumption that the $Z_t$ are i.i.d.
A quasi-posterior based on criterion (ref) is then obtained as
where $\pi(\theta,\mu) = \pi(\theta|\mu)\pi(\mu)$ is the joint prior over $(\theta,\mu)$ and $\Xi$ is the corresponding joint prior support. This quasi-posterior is constructed directly from the sample criterion function together with the prior on $(\theta,\mu)$, and therefore remains well defined even when $(\theta,\mu)$ is not point identified. In many applications, researchers may choose a flat prior over $\theta$, i.e., $\pi(\theta|\mu) \propto 1$. At the same time, the quasi-Bayesian framework also provides a convenient way to incorporate economically motivated prior information about $\theta$ when appropriate. In general, $p_{T}(\theta,\mu)$ will not be available analytically but will need to be approximated using Markov Chain Monte Carlo (MCMC) or other approximation methods; see, e.g., MCMCbook for a classic textbook introduction. We also provide a simple Gaussian approximation to the posterior in a setting where the prior for $\mu$ is taken to be normal with a small variance in Section (ref).
We obtain marginal posteriors for $\theta$ and $\mu$ as usual by integration:
where $\mathcal{M}$ and $\Theta$ respectively denote the support of $\mu$ and $\theta$. $p_T(\theta)$ captures posterior information about the economically meaningful parameter $\theta$ and thus is the chief object of interest. $p_T(\mu)$ also potentially provides interesting information by summarizing posterior beliefs about the plausibility term $\mu$.\footnote{Note that, while $\theta$ and $\mu$ are not jointly identified, the imposition of a proper prior over either $\theta$ or $\mu$ will lead to posterior updating over both $\theta$ and $\mu$, even in settings with $q = k$.}
Given the quasi-posterior, we can also immediately formulate optimal decisions under the quasi-posterior by minimizing the quasi-posterior expected risk. Specifically, let $\ell(\theta,\mu,d)$ be a loss function that depends on the underlying model parameters $(\theta,\mu)$ and a decision $d \in \mathcal{D}$.\footnote{In most applications, this loss function will depend only on the economically motivated parameter $\theta$, but we allow the loss to be over $\mu$ as well.} We can then define the expected loss minimizing decision under the quasi-posterior, denoted $s_T(p_T)$, as
The quasi-posterior (ref) and inferential objects obtained from it, such as optimal decisions or credible intervals, can be given an approximate Bayesian interpretation. fs:bayesmoment and andrews2022optimal study the classic semiparametric moment condition model under correct specification such that $m(\theta_*) \equiv 0$, i.e., under the dogmatic prior that $\mu \equiv 0$. fs:bayesmoment provide prior choices for the unknown model density such that the analog of (ref) within this setting corresponds to the posterior for $\theta$ as the limit when the prior becomes diffuse. By augmenting the parameter space to include $\mu$, (ref) can be obtained as the posterior over $(\theta,\mu)$ under the prior structure of fs:bayesmoment. andrews2022optimal study optimal decision rules in weakly identified moment condition models under correct specification such that $m(\theta_*) \equiv 0$. andrews2022optimal establish that the analog of (ref) for their setting, corresponding to (ref) under the dogmatic prior that $\mu \equiv 0$, results as the limit of a sequence of posteriors under a specific choice of proper priors. The resulting quasi-Bayes decision rule then corresponds to the pointwise limit of the sequence of Bayes decision rules and can thus be motivated as approximating the optimal Bayes decision. We show that these results continue to apply in our setting with non-dogmatic prior over $\mu$ in Section (ref). Of course, given the non-degenerate prior over $\mu$, the optimal Bayes decision depends explicitly on not only the prior for $\theta$ as in andrews2022optimal but also on the prior for $\mu$. See also kim:lilgmm and gallant:inducedspace for additional Bayesian motivation and perspective.
From a purely frequentist perspective, chernozhukov2003mcmc verify that inference based on the quasi-posterior is asymptotically equivalent to inference based on efficient GMM in strongly identified settings under correct specification ($m(\theta_*) \equiv 0$). chernozhukov2003mcmc further argue that basing frequentist estimation and inference on posterior summaries from (ref), such as using the posterior mean as a point estimator and posterior credibility interval as a confidence interval, may be desirable in settings where (ref) is hard to optimize. However, credible intervals resulting from the quasi-posterior (ref) within the partially identified setting where $\mu$ follows a non-degenerate prior no longer generally deliver usual frequentist coverage guarantees, though they do still have an approximate Bayes interpretation; see moon2012bayesian and gustafson:book.\footnote{andrews2022optimal also show that quasi-posterior interval estimates generally do not provide correct frequentist coverage in weakly identified settings and suggest a procedure to obtain confidence sets with proper frequentist coverage. We choose to focus our coverage results on strongly identified settings while allowing for a non-dogmatic prior over $\mu$. In principle, the weak identification robust confidence set construction of andrews2022optimal could also be incorporated in the present setting.}
We extend the results from chernozhukov2003mcmc by studying the frequentist properties of the quasi-Bayes posterior within our setting with a non-dogmatic prior over $\mu$ allowing for sequences with increasing $k$ and $q$ in Section (ref). Within this setting, we provide new Bernstein-von Mises type posterior concentration results showing that the quasi-posterior converges to a mixture of Gaussian distributions where the mixture weights and components depend heavily on the specific prior $\pi(\mu)$. That is, the posterior aligns with our intuition about partial identification in that the prior for $\mu$ plays a key role in the shape of the posterior even in the limit. See also gustafson:book. In the case that one has a dogmatic prior over $\mu$---e.g., $\mu \equiv \mu_*$---our concentration result reproduces chernozhukov2003mcmc, in the sense that our posterior approximation collapses to a Gaussian random variable with center $\theta(\mu_*)$ and variance equal to the limiting variance of the efficient GMM estimator in this case.
Based on the posterior concentration results, we then have that usual Bayesian credible regions from the quasi-posterior (ref) have correct frequentist coverage within a two-stage sampling thought experiment where each repeated sample corresponds to drawing a value $\mu_*$ from a random variable with density $\pi(\mu)$ and then generating data such that $m(\theta(\mu_*)) = \mu_*$. Alternatively, one can view this notion of coverage as providing an ex ante coverage guarantee in a setting where a single value $\mu_*$ will be realized from a random variable with density $\pi(\mu)$.
Given that the two-stage sampling notion of coverage is non-standard, we consider a final approach to leveraging (ref) to provide a confidence set with a uniform frequentist coverage guarantee when the plausibility characteristic is taken to be some fixed vector, $\mu_0$, whose value is unknown, but where it is known that $\mu_0 \in C$ for some known compact set $C$. The basic idea is that we can use QBPs exactly as in chernozhukov2003mcmc to obtain point and interval estimates that would be asymptotically equivalent to efficient GMM for any fixed $\mu_* \in C$ under strong identification. Letting $CI(\mu_*,\alpha)$ be the resulting $(1-\alpha)\times 100\%$ credible interval, it then follows that $\cup_{\mu_* \in C} CI(\mu_*,\alpha)$ has coverage at least $(1-\alpha)\%$ for $\theta(\mu_0)$. This approach mimics, e.g., the union of confidence intervals approach from conley2012plausibly and the approach outlined in Remark 3.3 of armstrong2021sensitivity.
In this section, we outline an approach to doing approximate quasi-Bayes inference when the prior for the plausibility term $\mu$ is normal with a small variance. Specifically, suppose that our prior is
for some fixed $q$ dimensional vector $\mu_0$ and a fixed, full-rank $q \times q$ matrix $\Lambda$.\footnote{We could relax the restriction that $\Lambda$ is full-rank at the cost of a modest complication of notation.} A simple form of $\Lambda$ is a diagonal matrix with $\lambda_k$'s on the diagonal, where a small $\lambda_k$ indicates that there is little uncertainty about the plausibility of the $k$-th moment and larger values indicate higher uncertainty.
Intuitively, prior (ref) captures the case where misspecification is believed to be small but non-zero in the sense that we believe the moment conditions “almost” hold with $m(\theta_0) = \mu_0$. Considering a sequence of priors with variance of order $1/T$ means that prior uncertainty concentrates at the same rate as the sample moments, so neither will dominate as we consider large $T$ asymptotic approximations. Following the literature, we refer to (ref) as a local prior; see, e.g. conley2012plausibly and armstrong2021sensitivity.
We now provide an approximation to the quasi-posterior with $\pi(\mu)$ defined in (ref) and a flat prior over $\theta$. We assume $\mu_0$ is such that a solution $\theta(\mu_0)$ satisfying $m(\theta(\mu_0)) = \mu_0$ exists. For simplicity, we further set $\mu_0 = 0$. We assume strong identification in the sense that $m(\theta(\mu))=\mu$ has unique solution $\theta(\mu)$ for each $\mu$ and that the following linearization around $\theta_0 \equiv \theta(\mu_0)$ holds: $$ m(\theta(\mu)) = G \left(\theta(\mu)- \theta_0\right) + o\left( \| \theta(\mu) - \theta_0\|\right), $$ where $G=\partial \mathbb{E}(\widehat{m}(\theta))/\partial \theta |_{\theta = \theta_0}$ and $G^{\top}G$ has minimal eigenvalue bounded away from zero. We focus here on the strong identification case because it yields a transparent approximation that captures important forces shaping the quasi-posterior. Further, define the weighting matrix $\widehat{A}_{T,\theta}$ and its population counterpart $A_\theta$:
Under these conditions, we show in Section (ref) that $p_{T}(\theta)$ is approximately proportional to $$ \exp\left(-T\|\widehat m(\widehat \theta) + G (\theta - \widehat \theta) \|_{\widehat{A}_{T,\theta}}^{2}/2\right), $$ where the quasi-posterior mode $\widehat{\theta}$ is the GMM estimator obtained using weighting matrix $\widehat{A}_{T,\theta}$. That is, we can approximate the quasi-posterior for $\theta$ as
Note that the weighting matrix $A_{\theta_0}$ in the quasi-posterior is different from the standard efficient GMM weighting matrix $\Omega(\theta_0)^{-1}$.\footnote{As $\widehat{A}_{T,\theta}$ will converge to $A_{\theta_0}$ with $T\to \infty$, we provide intuition as if $A_{\theta_0}$ were known.} Specifically, $A_{\theta_0}$ reflects additional uncertainty brought by not having a fixed, known value for $\mu$. We do see that we recover the case of efficient GMM by letting $\Lambda \to 0$---i.e. by assuming that there is no uncertainty over the moment conditions.
The quasi-posterior has several interesting features. The center of the quasi-posterior, $\widehat\theta$, corresponds to the classical GMM estimator that uses the weighting matrix $A_{\theta_0}$ rather than the efficient weighting matrix $\Omega(\theta_0)^{-1}$. This centering is intuitive as $\Omega(\theta_0)$ captures only sampling variation in the moments but does not reflect the additional uncertainty arising from the plausibility of the moments. The weighting matrix $A_{\theta_0}$ incorporates both sources of uncertainty, intuitively placing the most weight on moments about which the researcher is most confident in the sense that combined sampling variability and plausibility uncertainty is lowest.
Looking at the quasi-posterior variance, there are two further noteworthy features. First, the variance matrix $V = (G^{\top}A_{\theta_0} G)^{-1} \geq (G^{\top} \Omega({\theta_0})^{-1} G)^{-1}$, where $(G^{\top}\Omega({\theta_0})^{-1} G)^{-1}$ is the usual asymptotic variance of the efficient GMM estimator. The variance matrix $V$ thus captures additional uncertainty, relative to efficient GMM, introduced by a lack of certainty over the validity of the moment restrictions. Second, the approximate sampling distribution of $\widehat\theta$ is $$ \sqrt{T}(\widehat{\theta} -\theta_0 )\rightarrow_d \mathcal{N}(0, \bar V), \quad \bar V = (G^{\top} A_{\theta_0} G)^{-1} G^{\top} A_{\theta_0} \Omega({\theta_0}) A_{\theta_0} G (G^{\top} A_{\theta_0} G )^{-1}, $$ where $\bar V \leq V$ because $A_{\theta_0} \Omega({\theta_0}) A_{\theta_0} \leq A_{\theta_0}$. Thus, the quasi-posterior variance is also larger than the asymptotic variance of the $A_{{\theta_0}}$-weighted GMM estimator. Again, this larger quasi-posterior variance arises because the sampling distribution of the $A_{{\theta_0}}$-weighted GMM estimator is obtained under the dogmatic belief that $\mu \equiv 0$ and thus does not reflect uncertainty about the validity of the moment restrictions outside of through reweighting the moment conditions.
To summarize, the approximation in (ref) provides a very simple avenue to obtain approximate Bayesian inference under local Gaussian priors. While restrictive, it does seem like a Gaussian prior with small variance may provide a reasonable model for subjective beliefs about moment condition violations in some settings, and we illustrate the use of the approximation, along with illustrating simulation of the full quasi-posterior, in the empirical examples in the next section. More importantly, the approximation captures the clear intuition that there is no free lunch. Incorporating non-dogmatic priors over moment condition violations naturally results in less informative inference relative to the case where dogmatic priors are imposed---reflecting the researcher's uncertainty about the validity of the moment restrictions. This property seems desirable as the resulting inference likely more accurately reflects what can be learned in real empirical settings where model uncertainty exists.
This section applies plausible GMM in two illustrative empirical applications. In the first, we revisit acemoglu2001colonial, which uses linear instrumental variables (IV) to study the effect of institutions on economic output in a relatively small sample. In the second, we revisit CH:401krestat, which uses IV quantile regression to examine the effects of 401(k) participation on measures of household assets.
We start by illustrating our methodology by revisiting the classic study of acemoglu2001colonial, which investigates the effect of institutions on economic performance. The outcome variable $Y_t$ is the log of PPP-adjusted GDP per capita in 1995 where $t = 1, \ldots, 64$ indexes a set of countries that are ex-European colonies. The main regressor of interest, $X_t$, is a ten-point index measuring protection against expropriation risk, serving as a proxy for institutional quality. We consider a baseline specification from acemoglu2001colonial which includes normalized distance from the equator, $W_t$, as a control for geographic factors.
To address concerns about the endogeneity of $X_t$, acemoglu2001colonial adopt an IV strategy. Following acemoglu2001colonial, we consider two IV specifications. The first, denoted Linear IV(1), uses the log of settler mortality as the sole instrument. This specification corresponds to the baseline in the original study. The second specification adds the proportion of the population of European descent in 1900 as an additional instrument as is done in a robustness exercise in the original paper. This specification allows us to illustrate our procedure in a setting with an overidentifying moment restriction. We refer to this specification as Linear IV(2).
Formally, we consider the linear IV model $$ Y_t = \alpha + \beta_X X_t + \beta_W W_t + U_t, $$ with parameter vector $\theta = (\alpha, \beta_X, \beta_W)^\top$ and moment condition $$ g(Z_t, \theta) = (1, W_t, D_t^\top)^\top \left( Y_t - \alpha - X_t \beta_X - W_t \beta_W \right), $$ where $D_t$ denotes the vector of instruments.
To implement PGMM, we must specify priors for the parameters $(\theta, \mu)$. We set the prior for $\theta$ as $\mathcal{N}(0, \text{diag}(100, 4, 64))$. We set the prior variances for the elements of $\theta$ via a loose argument based on economic intuition. For example, we know that the $X_t$ is measured on a 10-point scale with empirical 25th and 75th percentiles equal to 5.6 and 7.8, respectively, and an empirical range of $3.5$ to $10$. A coefficient of $\beta_X$ of approximately $.5$ would thus suggest moving from the 25th to 75th percentiles of $X_t$ is associated with around a one log unit change (around a 170% change) in GDP, which seems economically quite large. We thus feel comfortable placing a relatively low prior probability on $\beta_X$ having a magnitude larger than 4. We use the same rationale for our choice of the prior over $\alpha$ and $\beta_W$.
To specify the prior for $\mu$, we use reasoning based on the IV model. Specifically, we consider a benchmark where model misspecification arises from an omitted latent factor, $C_t$, that is correlated with the exogenous variables and may directly affect the outcome. That is, we consider an “augmented” model $$ Y_t = \alpha + \beta_X X_t + \beta_W W_t + C_t + U_t, $$ where we represent $C_t$ as $C_t=(1, W_t, D_t^\top)\pi$ to capture misspecification that is linearly related to the exogenous variables. In the case of Linear IV(1), this augmented structure implies that, for a given $\pi$, the moment equation becomes $\mathbb{E}[g(Z_t, \theta)] = \mathbb{E}[(1, W_t, D_t)^\top (1, W_t, D_t)] \pi = \mu$.
We can then capture subjective beliefs about misspecification by placing a proper prior over $\pi$. We start by centering this prior at 0, reflecting the subjective belief that the arguments for exclusion and exogeneity are compelling enough to center our beliefs on no direct effect of the instrument and no endogeneity in the control variable. We then focus on the element of $\pi$ associated with the excluded instrument. Given that $D_t$ is log mortality among Europeans several hundred years prior to 1995, one might reasonably believe there is limited scope for the direct effect of $D_t$ to be large. We benchmark our prior using the subjective belief that, with high probability, the elasticity of GDP with respect to settler mortality is no larger than 10%, corresponding to the last entry of $\pi$ being no larger than 0.1. We encode these beliefs by specifying a mean-zero Gaussian prior for the entry of $\pi$ corresponding to $D_t$ with standard deviation $0.05$. We regard this as a sensible benchmark in this setting. It places substantial prior mass near 0, reflecting the view that the argument in acemoglu2001colonial is broadly persuasive, while still allowing for moderate violations and assigning less mass to larger departures. We complete the prior by assuming the other entries of $\pi$ behave similarly to the one corresponding to the instrument. Specifically, our baseline (denoted “PGMM-g”) uses a Gaussian prior given by $$ \mu \sim \mathcal{N}(0, \Sigma_T \Omega_d \Sigma_T^\top), $$ where $\Sigma_T = T^{-1} \sum_{t=1}^{T} (1, W_t, D_t^\top)^\top (1, W_t, D_t^\top)$ and $\Omega_d = 0.05^2 I_3$.
For Linear IV(2), we follow a similar approach. Here, the additional instrument is the proportion of the population of European descent. Assuming that its direct impact on the outcome is, with high probability, no greater than 0.01 (a semi-elasticity of 1%), we extend the Gaussian prior construction used in Linear IV(1) by setting $\Sigma_T = T^{-1} \sum_{t=1}^{T} (1, W_t, D_t^\top)^\top (1, W_t, D_t^\top)$ and $$ \Omega_d = \text{diag}(0.05^2 I_3, 0.005^2), $$ where the final diagonal entry corresponds to the new instrument.
Of course, it is important to gauge the sensitivity of the posterior to the prior specification. We thus report results using two additional simple prior settings. In the first, we consider a more diffuse prior for $\mu$ (denoted “PGMM(d)-g”), given by $\mathcal{N}(0, c \Sigma_T \Omega_d \Sigma_T^\top)$ with $c=4$. As a second alternative, we also consider a uniform prior for $\mu$ (denoted “PGMM-u”), distributed uniformly over the elliptical region $\mathcal{C} = \{ \left(\Sigma_T \Omega_d \Sigma_T^\top\right)^{1/2}c: c^{\top}c \leq \chi^2_{0.68}(q)\},$ where $q$ is the number of moment conditions, and $\chi^2_{0.68}(q)$ denotes the 0.68 quantile of $\chi^2(q)$. This prior thus allocates all probability mass to the 68% highest density region of the Gaussian prior used in the “PGMM-g” case.
We report the PGMM quasi-posterior obtained under our different priors, along with the quasi-posterior from chernozhukov2003mcmc obtained under the dogmatic prior $\mu \equiv 0$ (labeled "CH"), for $\beta_X$ in the specification with one excluded instrument in the top panel of Figure (ref).\footnote{Quasi-posteriors in the Linear IV(2) case, shown in Figure (ref), exhibit similar patterns. We also present posteriors for elements of $\mu$, which roughly align with the corresponding priors, in Figure (ref).} We see that, in terms of $\beta_X$, the quasi-posteriors are relatively robust to the prior over $\mu$. As anticipated, we see that the quasi-posteriors become somewhat more diffuse as the prior dispersion increases from $\mu \equiv 0$ to PGMM-g to PGMM(d)-g, although the changes in dispersion are relatively small despite the large increase in prior dispersion for $\mu$ across these cases. Unsurprisingly given the design of the uniform prior, we also see that both the benchmark Gaussian prior (PGMM-g) and related uniform prior (PGMM-u) produce very similar quasi-posteriors for $\beta_X$.
For additional insight, we provide 95% HPD intervals for $\beta_X$ under $\mu \sim \mathcal{N}(0, c \Sigma_T \Omega_d \Sigma_T^\top)$ for additional values of $c$ in the bottom panel of Figure (ref). Here, we see that the lower bound of the HPD interval is relatively stable for values of $c \leq 4$. However, the lower bound then decreases relatively quickly as $c$ increases away from 4, crossing 0 at $c \approx 4.5$. That is, under our prior for $\theta$ and class of priors for $\mu$, posterior mass remains largely concentrated over positive effects even allowing for what appear to us be economically large deviations from correct specification.
Finally, we observe that the left tails of the quasi-posteriors in Figure (ref) are very similar to the upper tail of the maintained prior for $\beta_X$. That is, it appears that the behavior of the upper tail of the posteriors may be driven largely by the prior choice for $\theta$. Consequently, the upper bounds of the provided intervals are relatively insensitive to the prior for $\mu$. While unsurprising, we find this interplay between prior structure interesting, especially as researchers often have reasonable economic understanding about plausible values for structural parameters.
We report interval estimates obtained from a variety of procedures under both the Linear IV(1) and Linear IV(2) specification in Table (ref). For frequentist methods, we report 95% level confidence intervals, and we report 95% level HPD regions for (quasi-)Bayesian procedures. Specifically, we consider intervals produced by applying the following:
All approaches allow for heteroskedasticity. 2SLS, CUE, CH, and S maintain the assumption of correct specification. All other procedures allow for departures from correct specification by relaxing the constraint that $\mu \equiv 0$, with AK being frequentist valid and the remaining procedures having a (quasi-)Bayes interpretation. We note that S is formally valid under weak identification, while formal frequentist results for the other procedures are obtained assuming strong identification.
Table (ref) shows that, with the exception of the AK interval, the qualitative conclusions are largely consistent across methods: The centers and lower bounds of the intervals lie above zero, suggesting a positive effect of institutions on output. More specifically, the interval estimates broadly fall within two groups, with 2SLS, CUE, Local Approx, and Local Approx(d) in one group and CH, S, PGMM-u, PGMM-g, and PGMM(d)-g in the other.
Interestingly, the 2SLS, CUE, and local approximation intervals all rely on asymptotic approximations obtained under strong identification and are substantially narrower than the remaining intervals. One possible interpretation is that these approximations are less reliable in this application because of weak-instrument concerns (see, e.g., chernozhukov2008instrumental, kleibergen2025double). Relatively weak identification could also explain the lack of updating, relative to the prior on $\beta_X$, in the upper tail of the distribution. That is, while we specified what seemed like an economically diffuse prior, the data do not seem sufficiently informative to shift beliefs in the upper tail of the effect distribution. Of course, the more relevant issue from an economic perspective is the amount of posterior mass in the left tail below zero. By contrast, the CH and PGMM intervals are obtained directly from the (quasi-)posterior rather than from strong-identification asymptotic approximations. It is therefore interesting that these intervals are quite similar to the interval produced by the weak-identification-robust procedure. Although this similarity may be partly specific to this application, it suggests a potentially useful connection that may be worth exploring further.
Looking at the quasi-Bayes procedures specifically, recall that the CH method imposes the validity of the moment conditions, while the PGMM procedures relax this assumption. This relaxation, of course, results in the PGMM intervals being wider than CH as the PGMM intervals reflect the added uncertainty from accounting for potential misspecification. However, at least under the priors considered, the increase in width is relatively small and does not qualitatively change the conclusions that one would draw relative to CH.
Finally, we observe that the AK intervals lead to qualitatively different conclusions than those from the other approaches. This difference is particularly pronounced in the Linear IV(1) specification, where the AK interval is substantially wider than the intervals produced by the alternative methods. The most informative comparison is between AK and PGMM-u, as both restrict the plausibility term to lie within the same support. The key distinction is that the AK approach is designed to ensure valid frequentist coverage by focusing on least favorable directions within the misspecified model, given only the support restriction on the plausibility term. In contrast, PGMM-u imposes proper subjective priors on both the plausibility term and the structural parameters, $\theta$. In this example, the subjective priors place very little mass on models with economically extreme values of $\beta_X$, resulting in the quasi-posterior assigning negligible mass to much of the AK interval. This outcome illustrates how subjective priors can substantially shape the quasi-posterior in partially identified settings. Rather than targeting worst-case combinations, the quasi-posteriors reflect economically motivated beliefs about the joint distribution of the structural parameters and the plausibility term. Both approaches serve meaningful purposes, but we believe there are scenarios in which inference based on the quasi-posterior under subjective priors may offer more economically relevant insights.
In this subsection, we illustrate the use of PGMM in a non-linear model by using IV quantile regression (IVQR) to estimate the impact of 401(k) participation on quantiles of net financial assets as in CH:401krestat. Specifically, we apply PGMM using the IVQR moment conditions from chernozhukov2005iv: $$ g_{\tau}(Z_t, \theta_{\tau})=(1, W_t^\top, D_t)^\top\left(\tau-{\mathbf{1}}{\left(Y_t \leqslant \alpha_\tau + X_t \beta_{X,\tau} + W_t^\top \beta_{W,\tau} \right)}\right) $$ where $\tau$ denotes the quantile of interest; $Y_t$ is the outcome variable representing net total financial assets (in 1991 dollars); $X_t$ is a binary indicator for 401(k) participation; $W_t$ is a vector of control variables; and $D_t$ is a binary instrument indicating 401(k) eligibility.\footnote{The covariates are income, a quadratic in age, family size, four indicators of education categories, marital status, two-earner status, defined benefit pension status, IRA participation, and home ownership. For more details about the data, see, e.g., abadie2003semiparametric and CH:401krestat.} We report results for three quantiles, $\tau \in \{0.15,0.5,0.85\}$, to illustrate performance for a low, central, and upper quantile.
The basic argument for 401(k) eligibility being a valid instrument for participation is that eligibility is determined by employers and so may plausibly be taken as exogenous after conditioning on job relevant covariates. See, e.g., abadie2003semiparametric for further discussion of the underlying exclusion restriction. Of course, there are reasons that one might worry that the exclusion restriction does not hold perfectly. For example, one might conjecture that firms that offered 401(k) plans were attractive to employees who prefer savings for other, unobserved, reasons. Motivated by such concerns, conley2012plausibly explore sensitivity of linear IV estimates of the effect of 401(k) participation on financial assets. Our analysis extends this line of work by investigating the sensitivity of quantile treatment effect estimates. We also note that examining quantile effects may be of substantive economic interest given the strong asymmetry of financial asset holdings.
Of course, implementing PGMM requires specification of priors for the structural parameters, $\theta_{\tau}$, and plausibility terms, $\mu_\tau$. In this example, we use the same diffuse prior for each $\theta_{\tau}$---$\theta_{\tau} \sim \mathcal{N}\left(0, \frac{10^{10}}{5} I_{14}\right)$---for all reported results. Given the magnitude of the outcome variable and units of the input variables,\footnote{For example, the 0.15 and 0.85 quantiles of $Y_t$ are approximately $-2751$ and $36,303$, respectively.} this prior seems to be extremely uninformative.
The more delicate choice is the prior over the local misspecification parameter $\mu_{\tau}$. As in the previous example, we assess sensitivity to this choice by considering both zero-mean Gaussian priors and zero-mean uniform priors for each of our three values of $\tau$. We set baseline priors using a stylized model for misspecification. Specifically, we construct a bound on the moment conditions, denoted $\delta_\tau$, under the assumption that any misspecification arises from a direct effect of $D_t$ on $Y_t$, capped at 2000 dollars in absolute value. To simulate this bound, we compute $$ \max_{\gamma \in \{-2000,2000\} } \left|\frac{1}{T}\sum_{t=1}^T (1, W_t^\top, D_t)^\top\left(\tau-{\mathbf{1}}{\left(\epsilon_t +D_t \gamma \leqslant 0 \right)}\right)\right|, $$ where the $\epsilon_t$ are generated as the residuals from the linear IV analog of our quantile models. Using the resulting $\delta_\tau$, we define priors for $\mu_{\tau}$ as $c \cdot \mathcal{N}\left(0, \text{diag}(\delta_\tau/3)^2\right)$ or as uniform priors over $[-c\delta_\tau/3, c\delta_\tau/3]$. To explore varying degrees of prior concentration, we consider $c \in \{0, 0.5, 0.9, 1.0\}$, where $c = 0$ corresponds to the dogmatic prior that maintains correct specification.
Figure (ref) displays marginal quasi-posteriors for both $\beta_{X,\tau}$ and the component of $\mu_\tau$ associated with the IV moment condition, denoted by $\mu_{D,\tau}$, under the Gaussian prior for $\mu_\tau$ with $c = 1$.\footnote{We provide quasi-posterior plots for the remaining settings in Figures (ref)-(ref).} For comparison, we also provide the marginal quasi-posterior for $\beta_{X,\tau}$ under the assumption of correct specification ($c = 0$) in dashed curves. We see that the quasi-posteriors for the quantile effect obtained under correct specification concentrate over positive values for each value of $\tau$, suggesting a robust positive effect of 401(k) participation on net financial assets. The quasi-posteriors under correct specification are suggestive of larger quantile treatment effects at higher quantiles.
Looking at the PGMM quasi-posteriors, we see that allowing for departures from correct specification according to our prior leads to substantially more diffuse quasi-posteriors. For $\tau = 0.15$ and $\tau = 0.85$, the quasi-posterior for $\beta_{X,\tau}$ now places substantial mass on both large positive and large negative values. This substantial spread in the posterior implies that it becomes difficult to draw reliable conclusions about the lower and upper quantile treatment effect of 401(k) participation once we allow for plausible violations of the exclusion restriction consistent with the instrument having up to a \$2,000 direct effect on savings. The resulting intervals for low and high quantiles from “PGMM-g” in Figure (ref) span from -7.88 to 18.09 and -13.66 to 33.23 (in units of thousands of dollars), respectively. In contrast, the quasi-posterior for the median remains concentrated over positive values, suggesting a relatively robust positive median treatment effect of 401(k) participation.
Interestingly, Figure (ref) reveals that the marginal quasi-posterior distributions of $\mu_{D,\tau}$ differ noticeably from the prior. In particular, the quasi-posterior for $\tau = 0.15$ ($\tau = 0.85$) is shifted to the left (right) relative to the prior. The quasi-posterior for the median remains centered approximately over the prior center, but has substantially thicker tails than the marginal prior. This pattern suggests that deviations from the baseline model are more likely at the lower and upper quantiles, indicating that the moment conditions for the tails are “less plausible”---in the sense that their quasi-posteriors are not centered at zero---than those for the median. We find it interesting that, at least viewed through the lens of the marginal quasi-posteriors over $\mu_{D,\tau}$, the combination of data and moments leads to updating in the direction of model misspecification. That is, the marginal quasi-posterior over the sensitivity term for the IV moment restriction is less concentrated around zero (correct specification) than the initial prior. We take this updating as further evidence that a researcher should be hesitant to trust results that dogmatically maintain correct specification in this example.
Figure (ref) reports 95% highest quasi-posterior density intervals constructed using the PGMM method under our full set of prior specifications for $\mu_\tau$. In addition to the PGMM intervals obtained from simulating the full quasi-posterior, we also report intervals based on the local limiting approximation described in Section (ref) and Theorem (ref), as well as frequentist intervals constructed using the method of armstrong2021sensitivity (AK). The AK intervals are derived under a local misspecification framework in which the true parameter value $\theta_{\tau,X}$ is assumed to satisfy $\sqrt{T} m(\theta_{\tau, X}) \in \mathcal{C}_\tau$. For comparability, we define the restriction set $\mathcal{C}_\tau$ to match the support of the corresponding uniform priors used in PGMM-u for a given constant $c$: $$ \mathcal{C}_\tau = \left\{ \sqrt{T} \cdot \text{diag}(c \delta_\tau / 3) a : a \in \mathbb{R}^q,\ |a|_\infty \leq 1 \right\}, $$ where $|\cdot|_\infty$ denotes the $\ell_\infty$ norm. The parameter $c$ in Figure (ref) thus has a different interpretation for the different methods. For PGMM with a Gaussian prior on $\mu_\tau$ (PGMM-g), it determines the standard deviations of the prior distribution. For PGMM with a uniform prior (PGMM-u), it sets the upper and lower bounds of the uniform prior support. For the local approximation intervals (Local Approx), it indexes the scale of the local Gaussian prior $\mathcal{N}(0, \Lambda_c / T)$, where $\Lambda_c = T c^2 \cdot \text{diag}((\delta_\tau / 3)^2)$, consistent with the PGMM-g case. For the AK intervals, $c \delta_\tau$ parameterizes the local misspecification set $\mathcal{C}_\tau$ as defined above.
As shown in Figure (ref), the results for the lower and upper quantiles appear relatively sensitive to both the value of $c$ and the method used to obtain the interval estimate. As expected, the AK intervals---which are designed to ensure asymptotic frequentist coverage under worst-case local misspecification---are strictly wider than PGMM-u intervals in all cases. For these lower and upper quantiles, we see that the AK intervals are much wider than the corresponding PGMM intervals and, interestingly, are tracked relatively closely by the intervals obtained from the local approximation to the quasi-posterior. An interesting feature of this example is that the moment corresponding to plausibility term $\mu_{D,\tau}$ has natural support restrictions. The full quasi-Bayes procedure updates such that values that violate these support restrictions have very little posterior mass. This updating does not occur in either the local approximation or the AK intervals, which may explain some of the discrepancy between the procedures, especially for larger values of $c$.
In contrast, the intervals for the median effect are notably more stable across methods. All approaches yield similar interval estimates, and the lower bounds remain above zero even under relatively diffuse priors on the misspecification term. This stability suggests that inference about the median treatment effect is more robust to the specification of priors and the choice of estimation method.
As in the previous example, we find that examining quasi-posteriors under non-dogmatic priors on the degree of misspecification offers valuable insight into the identification and plausibility of the estimated effects. Estimates of quantile effects in the upper and lower tails are relatively sensitive to assumptions about model specification, with this sensitivity manifesting as instability across methods and prior choices. In contrast, the estimated median effects appear considerably more robust, yielding qualitatively similar results across all considered specifications. Finally, the updating from the prior over $\mu_\tau$ to the quasi-posterior is particularly interesting. In all cases, the quasi-posteriors place more mass away from $\mu_\tau = 0$ than the prior, indicating quasi-posterior evidence against correct specification. This shift suggests that researchers should be cautious about imposing the assumption of correct specification too rigidly in this setting.
This section is organized as follows. We first define notation in Subsection (ref). Subsection (ref) introduces a Bayesian optimal decision-theoretic motivation for the proposed quasi-Bayes procedure. Subsection (ref) presents a Bernstein-von Mises (BvM) theorem in a fixed-dimensional setting under local misspecification and establishes theoretical coverage guarantees for the highest quasi-posterior region. Subsection (ref) extends the BvM result to the high-dimensional case. Finally, subsection (ref) provides a frequentist justification for the coverage of the Bayesian credible set.
For a vector $v=(v_1,\ldots,v_d)\in\mathbb{R}^d$ and $q>0$, we denote $|v|_q=\left(\sum_{i=1}^d|v_i|^q\right)^{1/q}$,{$|v|_{\infty}=\max_{1\le i\le d}|v_i|$}, and $\|v\|=|v|_2$. For a vector $v$ and a conformable non-negative definite matrix $A$, define $\|v\|_{A}:=\sqrt{v^{\top}Av} \geq 0$. For two positive number sequences $(a_T)$ and $(b_T)$, we say $a_T\lesssim b_T$ (resp. $a_T\asymp b_T$) if there exists $C>0$ such that $a_T/b_T\le C$ (resp. $1/C\le a_T/b_T\le C$) for all large $T$. We denote $a_T \ll b_T$ if $a_T/b_T\rightarrow 0$ as $T\rightarrow\infty$, and write $a_T\gg b_T$ if $a_T/b_T\rightarrow\infty$ as $T$ diverges. Let $\nu(\mu)$ denote a point sufficiently close to $\mu$ such that a unique solution $\theta(\nu(\mu))$ of $m(\theta(\nu(\mu))) = \nu(\mu)$ exists. Denote a total variation of moments (TVM) type norm of ${\kappa}$ for a real-valued measurable function $g$ on $\Theta\times \mathcal{M}$ by $\left\Vert g \right\Vert_{TVM(\kappa)} =\int_{\theta \in \Theta, \mu \in \mathcal{M}} (1+\|\sqrt{T}(\theta-\theta(\nu(\mu)))\|^{\kappa})|g(\theta,\mu)|d\theta d\mu$ for ${\kappa}\geq 0$.
We use $\propto$ to denote “proportional to”. We use the subscript $p$ to denote statements with respect to the outer measure $\mathbb{P}^*$ of a given probability $\mathbb{P}$. We use $\rightarrow_d$ to denote convergence in distribution. We set $(X_T)$ and $(Y_T)$ as two sequences of random variables. Write $X_T=O_{p}(Y_T)$ if $\forall \epsilon>0$, there exists $C>0$ such that $\mathbb{P}^*(|X_T/Y_T|\le C)>1-\epsilon$ for all large $T$. We denote $X_T=o_{p}(Y_T)$ if $X_T/Y_T\rightarrow_p 0$ as $T\rightarrow\infty$. We limit ourselves to situations in which, given $\mu$, observations are a random sample from a distribution $\mathbb{P}_{\mu}$ with $\mathbb{P}_{\mu}$ being the conditional law of the random sample given $\mu$; and probability statements under $\mathbb{P}$ are made relative to the joint distribution of the random sample and $\mu$, given a fixed latent distribution $F_\mu$ over $\mu$.
Given the quasi-posterior, we can formulate optimal decisions under the quasi-posterior by minimizing the quasi-posterior expected risk. Specifically, let ${\ell}(\theta,\mu,d)$ be a loss function that depends on the parameters $\theta, \mu$ and a decision $d\in\mathcal{D}$. Policymakers may want to choose a decision $d$ to minimize the loss ${\ell}\left(\theta,\mu,d\right)$ that depends on both $\theta$ and $\mu$. The expected loss-minimizing decision under the quasi-posterior, denoted by $s_T(p_T)$, then takes the usual form:
We note that this framework encompasses the setting where loss depends only upon $\theta$ as a special case.
These quasi-Bayes decision rules can be motivated as approximations to fully Bayesian rules. In the weak identification setting, andrews2022optimal show that the quasi-posterior based on continuously updated GMM can be obtained as the limit of a sequence of posteriors under proper priors. The corresponding quasi-Bayes decision rule can then be interpreted as the pointwise limit of the associated Bayes decision rules.
In the case where parameters are low-dimensional, the results of andrews2022optimal can readily be adapted to the PGMM framework. Specifically, we have that, under regularity conditions, e.g., Assumptions (ref) and (ref).ii) stated in Section (ref), the process $\sqrt{T}\widehat{m}(\cdot) - \sqrt{T} m(\cdot)$ converges in distribution to a mean-zero Gaussian process with covariance function $\Sigma(\cdot, \cdot)$ and mean function satisfying $m(\theta(\mu))=\mu$ on $\mu\in \Gamma$. It then follows that we can construct a likelihood as in andrews2022optimal by properly substituting their parameter $\theta^*$ with the pair $(\theta(\mu), \mu)$. As a result, the optimal quasi-Bayesian decision rule under model misspecification retains the desirable properties established by their analysis. We provide the supporting technical details in Supplementary Appendix Section (ref).
This subsection considers a local misspecification setting in which the prior on $\mu$ is Gaussian with variance $\Lambda/T$. We show that, under this specification, the quasi-posterior distribution is asymptotically Gaussian and coincides with the results in chernozhukov2003mcmc in the special case where $\Lambda = 0$---i.e., in the case that $\mu \equiv 0$. The formal result in this section, Theorem (ref), provides justification for the arguments presented in Section (ref).
We start by presenting the technical assumptions under which we establish the Gaussian approximation. Throughout this section, we set $\mu_0 = 0$ and define $$ G(\theta(\mu)) = \frac{\partial \mathbb{E}_{P_\mu}\left[g(Z_t, \theta(\mu))\right]}{\partial \theta(\mu)}. $$ We also let $\theta(\mu_0) = \theta_0$, $A_{\theta(\mu_0)} = A_{\theta_0}$, and $\Pi_T(\theta, \mu_0)$ be a density function with $$ \Pi_T(\theta, \mu_0) \propto \exp\left( -\frac{T}{2} \|\theta - \widehat{\theta}\|^2_{G(\theta_0)^\top A_{\theta_0} G(\theta_0)} \right). $$
Assumption (ref) defines the support of the plausibility characteristic and, importantly, a set $\Gamma$ within the support such that the moment condition has a unique solution for each value of $\mu \in \Gamma$. In Assumption (ref) below, we require that the prior for $\mu$ has positive mass over at least one point in $\Gamma$, which ensures the quasi-Bayes procedure is (asymptotically) well-behaved.
Assumption (ref) concerns the properties of the GMM estimator with a weighting matrix that incorporates prior uncertainty as outlined in Section (ref). It specifically imposes that the resulting GMM estimator has a linear expansion dominated by its leading term. This assumption rules out weak identification. Nonetheless, the main results, together with analogous insights about misspecification, could in principle be extended to settings with weak instruments along the lines of andrews2022optimal, albeit at the cost of substantial additional technical work. Assumption (ref) then characterizes the asymptotic behavior of the leading term in this expansion. Assumptions (ref) and (ref) are analogous to standard assumptions that align, for example, with Assumption 4 and conditions (ii) and (iii) in Proposition 1 in chernozhukov2003mcmc.
Assumption (ref) is a modulus-of-continuity condition similar to condition (iv) of Proposition 1 in chernozhukov2003mcmc, used there to handle non-smooth criterion functions. It requires the remainder term to be bounded in a neighborhood of $\theta_0$ and is satisfied when the moments are sufficiently smooth. The final condition in Assumption (ref) ensures that $\theta_0$ is asymptotically well-identified.
The next assumption imposes restrictions on the prior that are sufficient for verifying approximate normality of the quasi-posterior.
The main substantive restriction of Assumption (ref) is that the prior for $\mu$ is Gaussian with variance of the same order as sampling variation. We further impose that the prior variance is full rank and that, a priori, $\theta$ and $\mu$ are independent. The requirement that $\Lambda$ be of full rank can be relaxed.\footnote{For example, let $B\in\mathbb{R}^{q\times\tilde q}, \tilde q<q,$ and let $\Lambda_x\in\mathbb{R}^{\tilde q\times\tilde q}$ be full rank. Consider the prior for $\mu$ generated from $\mu = Bx, x\sim \mathcal{N}\bigl(0,\;T^{-1}\Lambda_x\bigr).$ Following the same arguments as used to establish Theorem (ref), we can establish the quasi-posterior density $p_T(\theta)$ is, for large $T$, approximately Gaussian with covariance matrix $A_\theta\;=\; \Omega(\theta)^{-1} \;-\; \Omega(\theta)^{-1}\,B\, \bigl(\Lambda_x^{-1}+B^\top\,\Omega(\theta)^{-1}B\bigr)^{-1} \,B^\top\,\Omega(\theta)^{-1}.$}
We now present the first main theorem, which shows that the quasi-posterior density $p_{T}(\theta)$ converges in the TVM norm to a Gaussian density. With slight abuse of notation, we denote $\bar{p}_T(\widehat{m}(\theta))=p_T(\theta)$ and obtain $\bar{p}_{T}\left(\widehat{m}(\widehat{\theta})-G(\theta_0)(\widehat{\theta}-\theta)\right)$ by replacing $\widehat{m}(\theta)$ in $\bar{p}_T(\widehat{m}(\theta))$ with $\widehat{m}(\widehat{\theta})-G(\theta_0)(\widehat{\theta}-\theta)$.
Theorem (ref) demonstrates that $p_T(\theta)$ can be asymptotically approximated by a Gaussian density function $\Pi_{T}(\theta, \mu_0)$ under a sequence of Gaussian priors over misspecification that concentrate at the same rate as sampling error. Further, the theorem confirms the expected result that $p_T(\theta)$ concentrates in a $1 / \sqrt{T}$ neighborhood of $\theta_0$ under local misspecification. This result differs from the related approximation result in chernozhukov2003mcmc in that the quasi-posterior depends on the prior for the plausibility characteristic, even asymptotically. We do note that the approximation result reproduces the result from chernozhukov2003mcmc under $\mu_0 =0$ and $\Lambda= 0$.
We note that Theorem (ref), along with Theorems (ref)-(ref) presented below, could be established without requiring the data stream $\{Z_t\}_{t=1}^{T}$ to be i.i.d. Rather, we could work with moments defined as $m_T(\theta(\mu)) = T^{-1}\sum_{t=1}^{T}\mathbb{E}_{\mathbb{P}_{\mu,t}}[g(Z_{t},\theta(\mu))] = \mu$ where $\mathbb{P}_{\mu,t}$ is the marginal distribution for $Z_t$. Within this structure, results could be established under suitable restrictions on dependence and heterogeneity. We do not pursue this direction formally to avoid further complicating notation.
This section discusses an extension of Theorem (ref) by allowing relatively general choice of prior for $\mu$. The key result is that, as in the previous section, the prior for $\mu$ matters even in the limit. However, under a general prior structure, we do not obtain a limiting Gaussian approximation. Rather, we have that, conditional on $\mu$, the limiting approximation is Gaussian. Thus, the limiting approximation to the posterior is a Gaussian mixture where mixture weights depend heavily on the prior. We establish the formal results under sequences that allow the dimensions $k$ and $q$ to grow with the sample size $T$, which offers a technical extension of some results even in the case where a dogmatic prior is placed over $\mu$.
To accommodate a broader family of weighting matrices, we allow the GMM-type criterion that serves to define our quasi-posterior, $Q_{T}(\theta,\mu)$, to be formed with any positive-definite weight matrix $\widehat W_{T}(\theta)$. That is, we now consider $$ Q_{T}(\theta,\mu)=- {T}\left(\widehat{m}(\theta)-\mu\right)^{\top}\widehat{W}_{T}(\theta)\left(\widehat{m}(\theta)-\mu\right), $$ where setting $\widehat W_{T}(\theta) = \widehat\Omega_{T}(\theta)^{-1}$ corresponds to the leading case discussed in previous sections. Within this more general formulation, we use $W(\theta)$ to denote the population analog of $\widehat{W}_{T}(\theta)$ in the same manner that $\Omega(\theta)$ serves as the population counterpart of $\widehat{\Omega}_{T}(\theta)$.
In stating the formal results in this section, we make use of additional notation. Assume that, for each $\mu \in \mathcal{M}$, there exists a unique $\nu(\mu)\in\Gamma$ such that \[ \|\mu-\nu(\mu)\| = \inf_{\gamma\in\Gamma}\|\mu-\gamma\|. \] Let $h(\theta,\mu) = G\big(\theta(\nu(\mu))\big) \big(\theta-\theta(\nu(\mu))\big) -\mu+\nu(\mu).$ For $\varepsilon\asymp C_\epsilon q\log T$, define the $\varepsilon/\sqrt{T}$-expansion of $\{(\theta(\mu),\mu):\mu\in\Gamma\}$ by \[
\] Thus, $B_{\varepsilon}$ is an $\varepsilon/\sqrt{T}$-neighborhood containing pairs $(\theta,\mu)$ that are close to pairs $(\theta(\mu),\mu)$ with $\mu\in\Gamma$. Let \[
\] where we abbreviate $h(\theta,\mu)$ as $h$, and \[
\] We now state sufficient conditions for establishing our limiting approximation to the quasi-posterior.
Assumption (ref) imposes regularity conditions that are sufficient for good behavior of the limiting criterion function. Importantly, these conditions ensure that, at fixed values of $\mu$, $\theta$ is strongly identified through the restrictions on $G(\theta)$; see, e.g., hansen2010instrumental.
The next assumption restricts priors, importantly requiring that $\theta$ and $\mu$ are a priori independent and that the marginal priors place positive mass uniformly over the corresponding parameter space. Define a neighborhood for $\mu$, \[ \mathcal{B}_\delta(\mu) := \left\{ \mu' \in \mathcal{M}: \| \mu' - \mu \| < \delta \right\}, \] and similarly define a neighborhood $\mathcal{B}_\delta(\theta)$ for a point $\theta \in \Theta$.
When \(q\) and \(k\) are of fixed dimension, the above assumption imposes a strictly positive lower bound on \(\pi(\theta(\mu))\) and a positive lower bound on the probability mass that \(\pi(\mu)\) assigns to a value \(\mu \in \Gamma\). This condition allows the prior on \(\mu\) to be very informative---for example, it may reduce to a single point within \(\Gamma\)---whereas the prior on $\theta$ must remain sufficiently diffuse to satisfy the required lower bound.
Before stating the next assumption, define \[ R_{T}(\theta,\mu) :=\frac{1}{2}\bigl(Q_T(\theta,\mu) +V_T(h(\cdot),\theta,\mu)\bigr) +\log\pi(\theta), \] which measures the discrepancy between the posterior and its Gaussian approximation. This assumption collects high-level rate conditions, which we verify under more primitive conditions in Lemmas (ref) and (ref) under Assumptions (ref)--(ref).
Assumption (ref) i) is verified by Lemma (ref) under Assumptions (ref)--(ref). Assumption (ref) ii) is verified by Lemma (ref) under Assumptions (ref) and (ref).
The condition $$ \sup_{\theta,\mu \in B_{\varepsilon}} \frac{T|R_{T}(\theta,\mu)|}{\|\sqrt{T}h(\theta,\mu)\|^2 + k(\log T)^2} \to_p 0. $$ arises from needing to control a modulus of continuity. It ensures the oscillatory behavior of the empirical process $ R_T(\theta, \mu) $ is mild. While the condition is high level, Lemma (ref) shows that this condition is satisfied with differentiable moments. Assumption (ref) effectively imposes an identification requirement for large values of $ \theta $ and a smoothness condition for smaller values of $\theta$ on the term $\sup_{(\theta,\mu) \in B_{\varepsilon}^{c}} TR_T(\theta,\mu).$ The assumption resembles the finite-sample bound in Lemma A.16 of spokoiny2019accuracy, and it enables the derivation of a tail bound outside the ball $B_{\varepsilon}$ using a Gaussian integral argument.
The first condition is used in Lemma (ref) to control the moment terms of order $\kappa$, and reduces to $k^3(\log T)^8/T\to 0$ when $\kappa=2$. The second condition implies both $q^2(\log T)^2/T\to 0$ and $q^2/(kT)\to 0$ since $q\geq k\geq 1$, and is needed to control the weighting matrix interaction term in Lemma (ref). It also ensures $q^2(\log T)^2\to\infty$, which guarantees the tail decay $k^{\kappa/2}\exp(-cq^2(\log T)^2)\to 0$ required in Lemma (ref)(ii).
Now, define $$N_{T}(\theta,\mu)=\frac{\exp\left\{ -\frac{1}{2}[V_T(h(.),\theta,\mu)]\right\} \pi(\mu)}{\int_{\mu \in \mathcal{M}}\int_{\theta \in \Theta}\exp\left\{ -\frac{1}{2}[V_T(h(.),\theta,\mu)]\right\} \pi(\mu)d\theta d\mu}.$$ Under the stated assumptions, we obtain the following Bernstein-von Mises-type result establishing that the quasi-posterior is asymptotically approximated by $N_T(\theta,\mu)$.
The above theorem indicates that, conditional on fixed values of $ \mu $, the quasi-posterior distribution of $ \theta $ can be well approximated by a Gaussian distribution, thus facilitating practical inference via conditional sampling, as demonstrated in Theorems (ref)-(ref). In contrast to Theorem (ref), Theorem (ref) relaxes the prior specification on $\mu$ by not imposing a Gaussian prior, thereby extending the applicability of the result. While the joint limiting distribution $ N_T(\theta, \mu) $ is not Gaussian in general, it becomes Gaussian when conditioning on $ \mu $. This insight has practical implications: one may select a representative set of $ \mu $ values, compute the corresponding conditional quasi-posterior distributions of $ \theta $, and aggregate the highest quasi-posterior density regions. This approach mirrors the strategy employed by conley2012plausibly for constructing robust inference under partial identification.
Let $PR_T(\alpha)$ denote the $(1- \alpha)$% ($0<\alpha< 1$) highest quasi-posterior density region for $\theta$ obtained from the quasi-posterior $p_T(\theta)$.\footnote{For a positive constant $c$ and density $p_T(\theta)$, $PR_T(\alpha) = \{\theta \in \Theta: p_T(\theta)> c\}$ such that $\int_{PR_T(\alpha)} p_T(\theta) d \theta = 1- \alpha$.} Lemma (ref) shows that $PR_T(\alpha)$ asymptotically provides valid weighted average frequentist coverage in large samples if we envision a world where nature draws $\mu$ from the prior $\pi(\mu)$, in which case $\pi(\mu)$ coincides with the fixed latent distribution $F_\mu$ in the data generating mechanism.
In this section, we provide an approach to obtain regions for $\theta$ that deliver valid frequentist coverage guarantees under support restrictions over $\mu$. The regions are constructed by taking unions of quasi-posterior credible regions for fixed values of $\mu$. Frequentist validity of this approach relies on properties of quasi-posterior intervals obtained from the posterior distribution $p_{T}(\theta,\mu)$ established in Theorems (ref)-(ref) in this section.
In the following, we suppose that one is interested in a continuously differentiable function $\eta(\theta): \mathbb{R}^{k}\to \mathbb{R}$. To state our results, we define the following quantities at fixed, given values of $\mu$: $$J_{\Omega,W}\left(\theta(\mu) \right) = G(\theta(\mu))^{\top}W(\theta(\mu))\Omega(\theta(\mu)) W(\theta(\mu))G(\theta(\mu)),$$ $$U_{T}(\mu)={J}_{W}\left(\theta(\mu)\right)^{-1}{\Delta}_{T,W}\left(\theta(\mu)\right), \; {J}_{W}\left(\theta(\mu)\right) = G(\theta(\mu))^{\top}W(\theta(\mu))G(\theta(\mu)), $$ $$\Delta_{T,W}\left(\theta(\mu)\right)=G(\theta(\mu))^{\top}W(\theta(\mu))(\widehat{m}(\theta(\mu))-\mu).$$ We first introduce a high level assumption for an estimator of $\theta(\mu)$ which is defined at a fixed value of $\mu$. While we focus on quasi-Bayes estimators in this paper, we note that the estimator in this section can be relatively generic. For example, it could be a LTE conditional on a value of $\mu$, $\widehat{\theta}(\mu) =\mbox{argmin}_{d\in \mathcal{D}} \int_{\theta} \ell(\theta,d) p_T(\theta,\mu)/p_{T}(\mu)d\theta$, or a CUE $\widehat{\theta}(\mu) =\mbox{argmin}_{\theta \in \Theta} Q_T(\theta,\mu)$, among many others.
The expansion assumed in Assumption (ref) is readily justified for GMM estimators; see, e.g., Corollary (ref). The proof of Theorem (ref) implies that, for any $\mu \in \Gamma$, the CUE admits the same first-order linearization as the LTE with symmetric loss functions analyzed in chernozhukov2003mcmc, when considering the quasi-posterior conditional on $\mu$; see also Theorem 2 of chernozhukov2003mcmc for the fixed-$k$ case. When the weight matrix satisfies the generalized information equality (Equation (ref)), Assumption (ref) further implies that the asymptotic variance of the leading term $(\partial \eta(\theta(\mu))/\partial\theta)^{\top} U_{T}(\mu)$ in the expansion of $\eta(\widehat{\theta}(\mu)) - \eta(\theta(\mu))$ is given by $$ T^{-1} \left(\frac{\partial \eta(\theta(\mu))}{\partial\theta}\right)^{\top} {J}_{\Omega}\left(\theta(\mu)\right)^{-1} \left(\frac{\partial \eta(\theta(\mu))}{\partial\theta}\right), \; {J}_{\Omega}\left(\theta(\mu)\right) = G(\theta(\mu))^{\top}\Omega(\theta(\mu))^{-1} G(\theta(\mu)). $$
We now state two theorems that make use of different features of quasi-posteriors obtained conditional on fixed values of $\mu$ to produce interval estimates for $\eta(\theta(\mu))$. The first main result in each theorem verifies that the resulting interval estimates have asymptotically correct frequentist coverage under the assumption that the fixed value of $\mu$ corresponds to the value of $\mu$ defining the conditional distribution from which data were realized. As a consequence, we can obtain frequentist confidence regions with correct coverage under the prior support condition that $\mu$ belongs to a known set $\mathcal{M}$ without requiring a completely specified prior by taking a union of confidence intervals produced at each $\mu \in \mathcal{M}$. This approach is analogous to the union of confidence intervals approach in conley2012plausibly and the approach outlined in Remark 3.3 of armstrong2021sensitivity.
Theorem (ref) shows that quantiles induced by the limiting conditional posterior density $\frac{N_T(\theta,\mu)}{N_T(\mu)}$ yield an asymptotically valid frequentist approximation to the distribution of $\sqrt{T}\big(\eta(\widehat{\theta}(\mu))-\eta(\theta(\mu))\big)$. The coverage result is driven by the generalized information equality in (ref), and is in line with existing results for point-identified scalar parameters and partially identified models; see, for example, chernozhukov2003mcmc and chen2018monte.
The first main result of Theorem (ref), equation (ref), verifies that posterior quantiles obtained from the quasi-posterior constructed conditional on fixed value of $\mu$ have asymptotically correct coverage under the corresponding conditional measure. Equation (ref), showing valid coverage of the union of intervals obtained under a support restriction, then immediately follows under the assumption that the value of $\mu$ under which the data were generated belongs to the specified support $\mathcal{M}.$ The results of Theorem (ref) critically depend on choosing $\widehat{W}_T(\theta(\mu))$ such that (ref) holds. Motivated by Theorem 4 of chernozhukov2003mcmc, the next result proposes an alternative procedure for using features of the quasi-posterior to construct intervals with frequentist coverage guarantees in settings where $\widehat{W}_T(\theta(\mu))$ is specified in such a way that (ref) does not hold.
Theorem (ref) verifies that intervals constructed making use of a normal approximation constructed conditional on fixed value of $\mu$ also have asymptotically correct coverage under the corresponding conditional measure. In practice, $\widehat{J}_{T}\left(\widehat{\theta}(\mu)\right)^{-1}$ in Theorem (ref) can be computed by multiplying the variance-covariance matrix of the MCMC sequence $S=\left(\theta^{(1)},\theta^{(2)},\ldots,\theta^{(B)}\right)$, where $B$ denotes the simulation sample size, by $T$ in settings where MCMC is used to approximate the quasi-posterior at a fixed value of $\mu$. As with Theorem (ref), it is then immediate that a union of intervals obtained using this approach under a support restriction on $\mu$ delivers valid frequentist inference.
Note that we state these results for completeness and to verify that the quasi-Bayes approach can be used to deliver frequentist valid inference under only support restrictions as in armstrong2021sensitivity. However, our chief interest is in using quasi-Bayes approaches in settings where informative prior information is of use. If valid frequentist inference under a support restriction is the goal, it is not clear there is much advantage to adopting the framework presented in this paper. Nevertheless, our posterior density assigns probability mass to both $\theta$ and $\mu$, enabling a ranking across different parameter values. This offers richer information than a simple interval estimate.
In this paper, we introduce Plausible GMM (PGMM), a quasi-Bayesian framework for inference in moment condition models that allows for potential misspecification. By placing a proper prior over the degree of misspecification, PGMM provides a flexible and transparent way to incorporate researchers' subjective beliefs about the plausibility of structural assumptions. This approach extends classical GMM by acknowledging that moment conditions are often credible but not exact, enabling more credible inference in the presence of model uncertainty.
Our theoretical contributions include posterior concentration results and new Bernstein-von Mises type approximations under partial identification for quasi-Bayes procedures allowing diverging dimensions of parameters and moments. In addition, we also provide decision-theoretic guarantees under some special cases. While not our main goal, we also provide an approach and results for using quasi-posteriors to obtain asymptotically valid frequentist inferential statements under support restrictions for the degree of misspecification.
Empirical applications illustrate the use of PGMM. In these examples, we see that PGMM intervals remain informative while allowing for subjective, but empirically motivated deviations away from dogmatic identifying assumptions. PGMM may thus offer a useful tool for applied researchers who wish to retain the structure of moment-based models while explicitly allowing for uncertainty about moments being perfectly satisfied.