Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
68,311 characters · 31 sections · 68 citation commands
Measurement Error and Counterfactuals in Quantitative Trade and Spatial Models
Economists use quantitative trade and spatial models to evaluate counterfactual scenarios. For instance, how do expenditure patterns across countries adjust in response to the implementation of a trade agreement? How are welfare levels affected when transportation infrastructure connecting regions is improved? These counterfactual questions are typically posed relative to an observed factual situation. This implies that the estimand of interest depends directly on the realized data, rather than on the underlying data-generating distribution—a departure from standard statistical settings.
The setting with data-dependent counterfactual estimands is further complicated by the fact that data in quantitative trade and spatial models are often measured with error goes2023reliability,linsi2023problem,teti2023missing. Unlike classical measurement error settings, where the estimand is typically a parameter of the correctly measured population distribution, here it is a functional of the realized data. To illustrate, consider the canonical Armington model armington1969theory, where predicted welfare changes from hypothetical trade cost shocks can be written as a function of baseline bilateral trade flows and the trade elasticity arkolakis2012new. The question I address is how measurement error in the observed trade flows affects uncertainty in the welfare predictions.
I develop an empirical Bayes framework for quantifying uncertainty around counterfactual predictions. The approach requires specifying both a measurement error model and a prior distribution over the latent true data, up to a set of hyperparameters. These hyperparameters are estimated from the observed data via an empirical Bayes step. Bayes’ rule then yields an estimated posterior distribution over the latent data given the noisy observations. Given the structure of quantitative trade and spatial models, this posterior induces a corresponding posterior over counterfactual predictions. Uncertainty can then be summarized by reporting an interval of posterior quantiles.
A natural accompanying point estimator for this procedure is the posterior median, which is always guaranteed to be contained in the reported interval. By contrast, the standard point estimator, which does not account for measurement error, may lay outside the interval. The posterior median answers the question: What does a Bayesian believe the counterfactual prediction would have been in the absence of measurement error? While this is a natural and intuitive question to ask, the answer necessarily depends on the prior.
In settings where the observed data consist of non-negative dyadic flows, I propose a default specification for the measurement error model and prior that can be calibrated directly from the data, yielding a widely applicable empirical Bayes approach. Specifically, I model measurement error as log-normal and use a log-normal prior centered on a structural gravity equation, with a point mass at zero to accommodate zero flows. This setup is designed for ease of implementation and is suitable for a wide class of quantitative trade and spatial models.
I consider two approaches to calibrating the hyperparameters in this default specification. The first assumes constant measurement error variances across flows and relies on researcher input or domain knowledge. The second, applicable in the case of international trade, uses the mirror trade dataset compiled by linsi2023problem, which reports bilateral trade flows as recorded by both exporters and importers. I interpret these paired observations as two independent noisy measurements of the true trade flow, enabling calibration of flow-specific measurement error variances.
To illustrate the impact of incorporating measurement error into counterfactual analysis, I revisit the applications in adao2017nonparametric and allen2022welfare. In adao2017nonparametric, which quantify the welfare effects of China’s accession to the WTO, I model measurement error in baseline bilateral trade flows. I apply the default empirical Bayes approach and construct uncertainty intervals that account for measurement error in the estimated changes in China’s welfare from 1996 to 2011. These intervals are substantially wider than those reported in adao2017nonparametric, which reflect estimation uncertainty.
In the setting of allen2022welfare, the counterfactual question concerns which highway links in the United States yield the highest return on investment and are therefore most promising for improvement. I model measurement error in traffic flows and apply the default empirical Bayes approach, calibrating the prior and measurement error model using estimates from musunuru2019applications. I compute uncertainty intervals that account for measurement error for the three links with the highest estimated returns. Although the intervals are wide, the relative ranking of the top three links remains robust.
This paper contributes to a growing body of work aimed at improving counterfactual analysis in quantitative trade and spatial models balistreri2008gravity,adao2017nonparametric,kehoe2017quantitative,adao2023putting,ansari2024quantifying,sanders2025new. The most closely related work is DingelTintelnot:2025, which studies calibration procedures in granular environments. That paper considers models that presume a continuum of agents and shows that, when data are limited, unit-level idiosyncrasies are absorbed into the model parameters, leading to overfitting and poor out-of-sample performance. My focus is on the complementary issue of uncertainty quantification due to measurement error—an issue that persists even in non-granular settings. DingelTintelnot:2025 recommends replacing raw observed data with fitted values from a low-dimensional model. I show how this recommendation can be nested into the proposed Bayesian framework.
The remainder of the paper is organized as follows. Section (ref) introduces the setting and notation. Section (ref) presents the empirical Bayes framework for accounting for measurement error in quantitative trade and spatial models. Section (ref) describes a widely applicable default approach. Section (ref) demonstrates the procedure in the context of the Armington model. Section (ref) applies the method to the trade setting in adao2017nonparametric and explores its use in the economic geography framework of allen2022welfare. Section (ref) concludes.
This section introduces the notation and discusses the key assumption that commonly underlies counterfactual analyses in quantitative trade and spatial models.
To begin, consider a baseline setting where there is no measurement error. Let $D\in\mathcal{D}\subseteq\mathbb{R}^{d_{D}}$ denote a data vector drawn from distribution $\mathcal{P}_{D}$, and let $\theta\in\Theta\subseteq\mathbb{R}^{d_{\theta}}$ denote a structural parameter. Our objective is to compute a scalar counterfactual quantity $\gamma\in\mathbb{R}$. The key assumption that the counterfactual object of interest has to satisfy is:
The exact functional form of $g$ depends on the specific quantitative model that is considered. In Appendix (ref) I discuss Assumption (ref) for two leading classes of models, namely invertible models and exact hat algebra models.
The main appeal of focusing on objects of the form in Assumption (ref) is that it allows researchers to answer counterfactual questions posed relative to a specific, observed factual situation. In quantitative trade and spatial settings, such questions are often at least as relevant as those concerning average effects. For example, in a quantitative model of international trade, the goal is typically to understand what would happen to the world following a specific policy change, rather than what would occur in a randomly drawn year under that policy.
Assumption (ref) implies that if the data $D$ are observed without error and the structural parameter $\theta$ is known, we can perfectly recover $\gamma$.\footnote{Indeed, by fixing $g$ I abstract away from model misspecification, an important problem I engage with in future work.} This contrasts with standard econometric models, where the object of interest is a function of the correctly measured distribution of the data, rather than the actual observations. So the key distinction with standard settings is:
Importantly, this difference implies that it would not suffice to be able to perfectly estimate the distribution $\mathcal{P}_{D}$. Towards uncertainty quantification, we hence need to account for uncertainty about the realized data themselves rather than their distribution.
This section introduces measurement error into quantitative trade and spatial models. It outlines how to quantify the resulting uncertainty for the counterfactual prediction of interest.
Under Assumption (ref), our object of interest can be written as a function solely of the data realizations and the structural parameter, which is convenient for answering relevant counterfactual questions. However, the data realizations are economic variables which are often measured with error. For instance, ortiz2018international and goes2023reliability highlight that there are large discrepancies between and within various data sources from trade and international economics. Motivated by this, I assume that, instead of the true data vector $D$, we observe a noisy version $\tilde{D}$.
For uncertainty quantification for the counterfactual prediction, we will require the posterior distribution of the true data given the noisy data. Towards that end, I introduce a model for the measurement error and a prior distribution for the true underlying data, \[
\] Here, $\vartheta\in\mathbb{R}^{d_{\vartheta}}$ is a vector of unknown hyperparameters.
Given such a prior distribution and a measurement error model, we can use an empirical Bayes approach to estimate the unknown hyperparameters.\footnote{Rather than estimating the parameters of the prior distribution for the true underlying data, which corresponds to an empirical Bayes approach, one could alternatively specify prior distributions for these parameters, which corresponds to a hierarchical Bayes approach.} Formally, we have \[ \tilde{\vartheta}=\underset{\vartheta}{\arg\max}\int\pi^{\mathrm{me}}\left(\tilde{D}|D,\vartheta\right)\pi^{\mathrm{prior}}\left(D|\vartheta\right)dD. \] Then, given the estimated hyperparameters $\tilde{\vartheta}$, we can use Bayes' rule to find the estimated posterior distribution of the true data given the noisy data,
Using this estimated posterior, we can generate draws for the true data given the noisy data.\footnote{Note that the measurement error distribution does not have to be mean zero, so also allows for measurement error bias. Nevertheless, even mean zero measurement error can result in bias in the counterfactual prediction of interest. This is automatically taken into account by the Bayesian approach when quantifying uncertainty. Furthermore, the individual measurement error distributions can be arbitrarily correlated in this general setup.}
The Bayesian approach allows researchers to incorporate economic knowledge through the prior. For example when considering measurement error in non-negative flows between locations, one can fit a prior centered on a gravity model, which I will do in Section (ref).
The object of interest is a function of the true data and the structural parameter. Going forward, I will assume the structural parameter is known, an assumption I will discuss in more detail in Section (ref). Then, under Assumption (ref) it follows that we can obtain the estimated posterior for the counterfactual object of interest, $\pi^{\mathrm{post}}\left(\gamma|\tilde{D},\tilde{\vartheta}\right)$.
Towards uncertainty quantification, we want to sample from this posterior and report the relevant quantiles.\footnote{Note that counterfactual predictions are typically derived as functions of the full system of counterfactual equilibrium variables. Thus, whether the researcher is ultimately interested in a scalar outcome, a relative comparison, or a global average, the mechanics of uncertainty quantification—drawing from the posterior over the true data and solving for equilibrium—remain the same.} The entire procedure is summarized in Algorithm (ref).
A natural accompanying point estimator for the procedure in Algorithm (ref) is the posterior median, which is always guaranteed to be contained in the reported interval. By contrast, the standard point estimator $g\left(\tilde{D},\theta\right)$, which does not account for measurement error, may lay outside the interval. The posterior median corresponds to an optimal estimate under the estimated posterior and under absolute value loss from a decision-theoretic perspective (see for example Proposition 2.5.5 in robert2007bayesian).
Note that the posterior median does not guarantee bias correction in the frequentist sense. It answers the question: What does a Bayesian believe the counterfactual prediction would have been in the absence of measurement error? While this is a natural and intuitive question to ask, the answer necessarily depends on the prior. But when the prior reflects well-established economic relationships—such as gravity patterns in quantitative trade and spatial models—the posterior median provides a principled estimate of the counterfactual prediction.
The literature on measurement error in nonlinear models is extensive, as reviewed in hu2015microeconomic and schennach2016recent, and the most closely related strand of measurement error literature is that on nonseparable error models matzkin2003nonparametric,chesher2003identification,hoderlein2007identification,matzkin2008identification,hu2008instrumental,schennach2012local,song2015estimating. However, these results do not apply to my setting.
The key distinguishing feature of the setting in this paper is that the object of interest $\gamma$ directly depends on the correctly measured data, because the equality in Assumption (ref) is an exact statement. This is convenient for answering counterfactual questions, and arises because counterfactual questions are typically posed relative to an observed factual situation. In contrast, in standard econometric methods of measurement error, the object of interest is a function of the correctly measured distribution of the data, $\mathcal{P}_{D}$, rather than the actual realized observations, $D$. This leads to the key distinction in Equation (ref).
This difference is important because in my setting, it would not suffice to be able to perfectly estimate the distribution $\mathcal{P}_{D}$. For example in a quantitative model of international trade, to answer counterfactual questions we need the realized trade flows, rather than the trade flow distribution from which they are drawn. In contrast, in standard econometric models of measurement error, knowing this distribution would suffice, because the estimands are functionals of the correctly measured distribution of the data. By virtue of that, we need to account for uncertainty about the observations themselves rather than their distribution.
The most relevant paper in the literature on improving counterfactual calculations in quantitative trade and spatial economics is DingelTintelnot:2025, which studies calibration procedures in granular settings. In these settings, individual idiosyncrasies do not wash out and can cause overfitting and poor performance out-of-sample. To deal with this, DingelTintelnot:2025 proposes to, instead of the observed data, either use fitted values obtained using a low-dimensional model or smooth the data using matrix approximation techniques. Both of these approaches can be cast as special cases of the proposed procedure in Algorithm (ref), by choosing a specific prior.
Specifically, the main recommendation is to use a low-dimensional model and is called the covariates-based approach. DingelTintelnot:2025 considers a quantitative spatial model with $L$ individuals. Let $\ell_{ij}$ denote the share of people residing in location $i$ and working in location $j$. The covariates-based approach then interprets the observed commuting shares $\left\{ \tilde{\ell}_{ij}\right\} $ as a finite sample from a continuum model. This results in the maximum likelihood model
for $\vartheta$ a set of hyperparameters and $h_{ij}\left(\vartheta\right)$ a model function which I discuss further in Appendix (ref).
The covariates-based approach in DingelTintelnot:2025 first finds a maximum likelihood estimator $\tilde{\vartheta}$ for $\vartheta$ using the model in Equation (ref). Next, focusing on a specific counterfactual object of interest denoted by $\gamma=g\left(\left\{ \ell_{ij}\right\} ,\theta\right)$ for some known structural parameter $\theta$ and function $g$, the approach recommends using the fitted values $\left\{ h_{ij}\left(\tilde{\vartheta}\right)\right\} $ instead of the observed shares $\left\{ \tilde{\ell}_{ij}\right\} $ to compute counterfactuals. That is, the main recommendation is to use the estimate \[ \tilde{\gamma}^{DT}=g\left(\left\{ h_{ij}\left(\tilde{\vartheta}\right)\right\} ,\theta\right) \] instead of $g\left(\left\{ \tilde{\ell}_{ij}\right\} ,\theta\right)$.
To see how the covariates-based approach from DingelTintelnot:2025 is nested in my Bayesian framework, consider the following prior and measurement error model,
where $\delta_{h_{ij}\left(\vartheta\right)}$ denotes the Dirac mass at $h_{ij}\left(\vartheta\right)$, implying a degenerate prior. The empirical Bayes step then combines the prior and measurement error model to find \[ \left\{ \tilde{\ell}_{ij}\cdot L\right\} |\vartheta\sim\mathrm{Multinomial}\left(\left\{ h_{ij}\left(\vartheta\right)\right\} \right), \] which overlaps with the model in Equation (ref), and uses maximum likelihood estimation to estimate $\vartheta$ by $\tilde{\vartheta}$. This yields the estimated prior and measurement error model \[
\] Using Bayes' rule we can then find the estimated posterior for the true commuting shares,
Note that after the empirical Bayes estimation step, since the estimated prior is degenerate, no information is taken from the estimated measurement error model.
The estimated posterior distributions in Equation (ref) translate to an estimated posterior for $\gamma$, \[ \pi^{\mathrm{post}}\left(\gamma|\left\{ \tilde{\ell}_{ij}\right\} ,\tilde{\vartheta}\right)=\delta_{g\left(\left\{ h_{ij}\left(\tilde{\vartheta}\right)\right\} ,\theta\right)}. \] Indeed, this posterior is a point mass at the counterfactual prediction that uses the fitted values $\left\{ h_{ij}\left(\tilde{\vartheta}\right)\right\} $ as inputs. It follows that the covariates-based approach is a special case of Algorithm (ref) by choosing the prior and measurement error model as in Equation (ref).
Note that if we follow Algorithm (ref) exactly, then in step 4a—where we draw from the posterior distributions in Equation (ref)—we will obtain the same values in each bootstrap iteration. As a result, the interval constructed in step 6 will collapse to a single point. This outcome is unsurprising, as the procedure in DingelTintelnot:2025 focuses on point estimation and correcting overfitting bias, rather than on quantifying estimation uncertainty. Accordingly, if the sole concern is overfitting bias, their method provides an appropriate approach.
The second recommendation in DingelTintelnot:2025 is to replace the observed data with a smoothed version using matrix approximation techniques. I discuss this approach in Appendix (ref).
The counterfactual prediction of interest will typically depend on a structural parameter $\theta$. It is common in applied work to plug in a fixed value for the structural parameter taken from the literature or obtained through data-driven methods, thus ignoring the uncertainty associated with the estimation process. An exception is adao2017nonparametric, which reports confidence sets for the counterfactual predictions of interest that account for estimation error.
Towards accounting for estimation error for quantitative trade and spatial models in the presence of measurement error, let $\tilde{\theta}$ denote the estimator of the estimand $\theta$. This estimand is usually a function of the distribution of the data $\mathcal{P}_{D}$. This implies that, to address measurement error affecting the structural parameter, one can apply the frequentist measurement error techniques discussed in Section (ref) to find a bias-corrected estimate, though the resulting correction will not admit a Bayesian interpretation.
Alternatively, in Appendix (ref) I outline a fully Bayesian approach that also considers estimation error. Specifically, I assume that the posterior distribution of the structural parameter $\theta$ given the true data $D$ is approximately normal, which is justified under regularity conditions that are closely related to those required for frequentist asymptotic normality. We then have two different posteriors, \[
\] As in Section (ref), a natural point estimator for the structural parameter is the median of the estimated posterior given the noisy data,
\[ \pi^{\mathrm{post}}\left(\theta|\tilde{D},\tilde{\vartheta}\right)=\int\pi^{\mathrm{post,ee}}\left(\theta|D\right)\pi^{\mathrm{post,me}}\left(D|\tilde{D},\tilde{\vartheta}\right)dD. \] Using Assumption (ref), we can also find the posterior $\pi^{\mathrm{post,ee}}\left(\gamma|D\right)$, and it follows that a natural point estimator for the counterfactual prediction is the median of the estimated posterior
In Appendix (ref) I describe how to sample from the estimated posteriors $\pi^{\mathrm{post}}\left(\theta|\tilde{D},\tilde{\vartheta}\right)$ and $\pi^{\mathrm{post}}\left(\gamma|\tilde{D},\tilde{\vartheta}\right)$. There, I also outline how to quantify uncertainty while jointly accounting for estimation error and measurement error in a natural way.
It is important to understand that Bayesian estimators such as the medians of $\pi^{\mathrm{post}}\left(\theta|\tilde{D},\tilde{\vartheta}\right)$ and $\pi^{\mathrm{post}}\left(\gamma|\tilde{D},\tilde{\vartheta}\right)$ need not satisfy frequentist properties such as consistency, even when the prior is well-specified. I elaborate on this possibility in Appendix (ref), by showing frequentist inconsistency of a structural estimator in a stylized example.
However, for models satisfying Assumption (ref), I am not aware of a frequentist framework that jointly incorporates estimation error and measurement error within a unified procedure that permits both point estimation and uncertainty quantification. By contrast, the proposed Bayesian approach accommodates both sources of uncertainty within a single coherent framework. Moreover, it nests the case in which only measurement error is present: as estimation error vanishes, the procedure naturally reduces to the framework that accounts solely for measurement error. I therefore recommend the Bayesian approach when both sources of uncertainty are relevant, and either the Bayesian or frequentist approach when only estimation error is of concern. In the latter case, the proposed framework can still serve as a robustness check with respect to measurement error.
This section proposes a default empirical Bayes approach that can be applied in many settings, as it is both economically reasonable for many quantitative trade and spatial models and computationally convenient. It also discusses the toolkit that accompanies the paper.
In many applications, the support of the true data and the sampling process naturally suggest both a measurement error model and a prior. For example, when observing migration shares constructed from count data, a multinomial measurement error model arises naturally from sampling variation. Since the true shares are non-negative and must sum to one, a Dirichlet prior provides a coherent and tractable specification.
For settings where such natural choices are not available, this section proposes a widely applicable default approach for quantifying uncertainty about the counterfactual prediction of interest. This default approach can be applied out-of-the-box to many quantitative trade and spatial models, but can also be easily adapted to other settings. It recommends default choices for the prior distribution and measurement error model, and discusses how to calibrate both based on observed data.
Concretely, consider the setting where we can write $\gamma=g\left(\left\{ F_{ij}\right\} ,\theta\right)$, for $\left\{ F_{ij}\right\} $ a set of non-negative flows between locations. This setup is commonplace in quantitative trade and spatial models costinot2014trade,redding2017quantitative,proost2019can. I assume that both the prior distributions on the true flows and the measurement errors are mixtures of a point mass at zero and a log-normal distribution, a so-called spike-and-slab distribution mitchell1988bayesian. The point mass at zero is necessary because in both trade and spatial applications bilateral flows of zeros are common, particularly when considering more granular data helpman2008estimating,DingelTintelnot:2025. This prior and measurement error model imply that the posterior distribution of the true flows given the noisy flows will also be a spike-and-slab distribution. This mixture model is fairly flexible, and conjugacy ensures computational tractability. Furthermore, I assume that the prior mean exhibits a gravity relationship, for which there is strong empirical evidence head2014gravity,allen201813.\footnote{One can easily enrich this gravity prior by adding other “distance” variables such as differences in income or productivity, or by adding dummies that indicate similarity such as contiguity or a common language, see for example silva2006log. I experimented with this but the results do not change much.} This is summarized in the following assumption:
The probability that a bilateral trade flow is truly zero is denoted by $p_{ij}$, and a true zero flow is assumed to always result in an observed zero.\footnote{Assumption (ref) implies that both true and spurious zeros occur randomly. Alternatively, one could think about incorporating endogenous zeros using selection mechanisms such as in helpman2008estimating.} The probability of a spurious zero—that is, an observed zero despite a non-zero underlying true flow—is denoted by $b_{ij}$. The prior means and variances are denoted by $\left\{ \mu_{ij}\right\} $ and $\left\{ s_{ij}^{2}\right\} $, respectively. The flow-specific measurement error variances are denoted by $\left\{ \varsigma_{ij}^{2}\right\} $.
Gather the hyperparameters in $\vartheta=\left(\left\{ p_{ij}\right\} ,\left\{ b_{ij}\right\} ,\beta,\left\{ \alpha_{i}^{\mathrm{orig}}\right\} ,\left\{ \alpha_{i}^{\mathrm{dest}}\right\} ,\left\{ s_{ij}^{2}\right\} ,\left\{ \varsigma_{ij}^{2}\right\} \right)$. It follows that the posterior distribution for the true flow between location $i$ and $j$, $F_{ij}$, given its noisy version, $\tilde{F}_{ij}$ is given by
for $i,j=1,...,n$, where $Q_{ij}\sim\mathrm{Bern}\left(\frac{p_{ij}}{p_{ij}+b_{ij}\left(1-p_{ij}\right)}\right)$.\footnote{Note that $F_{ij}$ are drawn independently across $ij$-pairs. As a result, certain adding-up constraints may no longer hold. In principle, this issue could be addressed by drawing from a constrained joint distribution.}
Conditional on being able to calibrate the parameters $\vartheta$—the empirical Bayes estimation step—one can quantify uncertainty about $\gamma$ by finding the interval as described in Algorithm (ref). Then, a default procedure for quantifying uncertainty about $\gamma$ is summarized in Algorithm (ref).
The hyperparameters in $\vartheta$ need to be calibrated. I consider two cases.
In the baseline case I restrict the measurement error variances and prior variances to be constant across flows so that $\varsigma_{ij}^{2}=\varsigma^{2}$ and $s_{ij}^{2}=s^{2}$ for all $i,j=1,...,n$. Furthermore, I require knowledge of the common measurement error variance $\varsigma^{2}$ and of the Bernoulli parameters $\left\{ p_{ij}\right\} $ and $\left\{ b_{ij}\right\} $.\footnote{In the absence of a prior on the measurement error variance, one could adopt a sensitivity analysis approach by plotting intervals while varying the variance. Alternatively, one could determine the minimum level of measurement error that would overturn the counterfactual conclusion.} It then remains to estimate $\left(\beta,\left\{ \alpha_{i}^{\mathrm{orig}}\right\} ,\left\{ \alpha_{i}^{\mathrm{dest}}\right\} ,s^{2}\right)$. Towards this, we can combine the equations in Assumption (ref) to find \[ \log\tilde{F}_{ij}\sim\mathcal{N}\left(\beta\log\mathrm{dist}_{ij}+\alpha_{i}^{\mathrm{orig}}+\alpha_{j}^{\mathrm{dest}},s^{2}+\varsigma^{2}\right),\quad\tilde{F}_{ij}>0. \] Using maximum likelihood estimation, it follows that the prior mean parameters can be estimated from the regression \[ \log\tilde{F}_{ij}=\beta\log\mathrm{dist}_{ij}+\alpha_{i}^{\mathrm{orig}}+\alpha_{j}^{\mathrm{dest}}+\phi_{ij},\quad\tilde{F}_{ij}>0, \] with $\phi_{ij}$ an error term. It follows that the estimated prior means and variance are
Obtaining estimators for these prior means and variances is what walters2024empirical calls the deconvolution step.
When the non-negative bilateral flows correspond to trade flows between countries, I use the mirror trade dataset from linsi2023problem to calibrate $\vartheta$. This dataset has two estimates of each bilateral trade flow, both as reported by the exporter and as by the importer. linsi2023problem shows that there are so-called mirror discrepancies in bilateral trade flows between almost all countries. This means that, for instance, while the value that Germany reports it imported from France and the value that France reports it exported to Germany should be the same, in practice they are often different. I interpret this as observing two independent noisy observations per time period for each bilateral trade flow. The key identifying assumptions are that the flow-specific probabilities of true zeros, the flow-specific probabilities of spurious zeros, and the flow-specific measurement error variances are constant over time.
The details for the calibration can be found in Appendix (ref). I first calibrate the probabilities of true zeros $\left\{ p_{ij}\right\} $ and the probabilities of spurious zeros $\left\{ b_{ij}\right\} $ by noting that for each bilateral trade flow we can use the time variation to identify the probabilities of observing a certain number of zeros. I then leverage the model structure to calibrate the measurement error variances $\left\{ \varsigma_{ij}^{2}\right\} $. Lastly, I calibrate the prior parameters, using a similar approach as for the baseline case with domain knowledge. To leverage country information and the fact that importers and exporters can differ in their reliability, I shrink the measurement error and prior variances using country-origin and country-destination fixed effects.
Accompanying the paper, I provide an easy-to-use toolkit that consists of three programs.\footnote{The toolkit is written in MATLAB and can be found on my website, https://sandersbas.github.io/. A version in R is available upon request.} The first program implements the high-level approach in Algorithm (ref). It takes as inputs $\left(B,\theta,\tilde{D},\pi^{\mathrm{post}}\left(D|\tilde{D},\tilde{\vartheta}\right),g\right)$ and outputs posterior draws $\left\{ \gamma_{b}\right\} _{b=1}^{B}$. The second program implements the default approach in Algorithm (ref). It takes as inputs $\left(B,\theta,\left\{ \tilde{F}_{ij}\right\} ,\tilde{\vartheta},g\right)$ and again outputs posterior draws $\left\{ \gamma_{b}\right\} _{b=1}^{B}$. The third program, which can serve as an input to the second, uses the mirror trade dataset of linsi2023problem and allows the researcher to choose countries and years for which they want to estimate the hyperparameters of the prior and measurement error model. This is summarized in Algorithm (ref).
This section illustrates the proposed procedure using the Armington model armington1969theory, a canonical workhorse model in international trade, as outlined, for example, in costinot2014trade.
Countries are indexed by $i,j=1,...,n$, and with CES preferences and perfect competition, it follows that the relevant gravity equations and budget constraints are:
Here, $F_{ij}$ denotes the trade flow from country $i$ to $j$, and $Y_{i}=\sum_{\ell=1}^{n}F_{i\ell}$, $E_{i}=\sum_{k=1}^{n}F_{ki}$ and $\kappa_{i}=\left(E_{i}-Y_{i}\right)/Y_{i}$ denote country $i$'s total income, total expenditure and the ratio of the trade deficit to income, respectively. Furthermore, $\tau_{ij}$ denotes the iceberg trade cost between country $i$ and $j$, which means that in order to sell one unit of a good in country $j$, country $i$ must ship $\tau_{ij}\geq1$ units, with $\tau_{ii}=1$. Lastly, $\varepsilon>0$ is the trade elasticity and $\left\{ \chi_{ij}\right\} $ are idiosyncratic preferences.
Now, say we are interested in the counterfactual where we change the trade costs $\left\{ \tau_{ij}\right\} $ proportionally by $\left\{ \tau_{ij}^{\mathrm{cf,prop}}\right\} $, holding the trade elasticity $\varepsilon$, the idiosyncratic preferences $\left\{ \chi_{ij}\right\} $ and the trade imbalance variables $\left\{ \kappa_{i}\right\} $ constant. In Appendix (ref) I show that we can then solve for the corresponding proportional changes in income, $\left\{ Y_{i}^{\mathrm{cf,prop}}\right\} $, using \[ Y_{i}^{\mathrm{cf,prop}}Y_{i}=\sum_{j}\frac{\left(\tau_{ij}^{\mathrm{cf,prop}}Y_{i}^{\mathrm{cf,prop}}\right)^{-\varepsilon}}{\sum_{k}\lambda_{kj}\left(\tau_{kj}^{\mathrm{cf,prop}}Y_{k}^{\mathrm{cf,prop}}\right)^{-\varepsilon}}\lambda_{ij}\left(1+\kappa_{j}\right)Y_{j}^{\mathrm{cf,prop}}Y_{j},\qquad i=1,...,n, \] where $\lambda_{ij}=F_{ij}/E_{j}$ denotes the expenditure share that country $j$ spends on goods from country $i$. By Walras' Law, the proportional changes in income are only pinned down up to a multiplicative constant. Subsequently, following costinot2014trade, we can exactly solve for proportional changes in expenditure shares and welfare (real consumption) levels:
The income levels $\left\{ Y_{i}\right\} $, the expenditure shares $\left\{ \lambda_{ij}\right\} $ and the trade deficit variables $\left\{ \kappa_{i}\right\} $ are all functions of the trade flows $\left\{ F_{ij}\right\} $, so the relevant counterfactual mapping is \[ \left\{ F_{ij}\right\} ,\left\{ \tau_{ij}^{\mathrm{cf,prop}}\right\} ,\varepsilon\mapsto\left\{ W_{i}^{\mathrm{cf,prop}}\right\} . \] It follows that for a given counterfactual question as described by $\left\{ \tau_{ij}^{\mathrm{cf,prop}}\right\} $, we only require knowledge of the baseline trade flows $\left\{ F_{ij}\right\} $ and the trade elasticity $\varepsilon$. So we have:
The specific counterfactual question I consider is a 10% increase in all bilateral trade costs between 76 countries, so that $\tau_{ij}^{\mathrm{cf,prop}}=1+0.1\cdot\mathbb{I}\left\{ i\neq j\right\} $ for $i,j=1,...,n$. I focus on the proportional percentage changes in welfare in the Central African Republic, the Netherlands, Sweden and the United States. It follows that, fixing $\left\{ \tau_{ij}^{\mathrm{cf,prop}}\right\} $, we have
for each $q\in\left\{ \mathrm{CAF},\mathrm{NLD},\mathrm{SWE},\mathrm{USA}\right\} $.
For the Armington model, I will consider measurement error in trade flows $\left\{ F_{ij}\right\} $. Hence, instead of the true trade flows we observe noisy trade flows $\left\{ \tilde{F}_{ij}\right\} $, which in turn lead to noisy counterfactual predictions $\tilde{\gamma}_{q}$ for $q\in\left\{ \mathrm{CAF},\mathrm{NLD},\mathrm{SWE},\mathrm{USA}\right\} $. Our goal is to quantify the uncertainty that arises from measurement error, and report an accompanying point estimator that aims to answer the question what a Bayesian believes the counterfactual predictions would have been in the absence of measurement error.
If we specify a prior $\pi^{\mathrm{prior}}\left(\left\{ F_{ij}\right\} |\vartheta\right)$ and a measurement error model $\pi^{\mathrm{me}}\left(\left\{ \tilde{F}_{ij}\right\} |\left\{ F_{ij}\right\} ,\vartheta\right)$, we can use empirical Bayes estimation and Bayes' rule to find the estimated posterior $\pi^{\mathrm{post}}\left(\left\{ F_{ij}\right\} |\left\{ \tilde{F}_{ij}\right\} ,\tilde{\vartheta}\right)$. The default approach from Section (ref) can be applied. For the empirical Bayes step, the calibration of $\vartheta$, we can use the mirror trade data setting from Section (ref). So we can use the provided toolkit to obtain draws from $\pi^{\mathrm{post}}\left(\left\{ F_{ij}\right\} |\left\{ \tilde{F}_{ij}\right\} ,\tilde{\vartheta}\right)$. I fix the trade elasticity to $\varepsilon=5$, a typical value in the literature, which is also used in costinot2014trade.
We can see the impact of measurement error in Table (ref) and Figure (ref). In Table (ref) I compare the standard point estimates based on noisy flows and the posterior median estimates, and report the intervals obtained using Algorithm (ref). In Figure (ref) I plot the standard point estimates and the smoothed estimated posterior distributions.
We observe that for the Central African Republic and the Netherlands there is a considerable difference between the point estimate and the median of the posterior distribution, causing the point estimate to lie outside the credible set. For Sweden and the United States there is less of a discrepancy. These plots illustrate that the proposed approach automatically incorporates bias induced by measurement error, and that this bias can be both negative and positive. In Appendix (ref), I show the results for all 76 countries in the sample.
In this section I discuss the applications in adao2017nonparametric and allen2022welfare. The application in adao2017nonparametric extends the prototypical example to a panel-data setting and recovers changes in trade costs using variation over time. The application in allen2022welfare illustrates how the proposed procedure can be extended to an economic geography framework.
The empirical application of adao2017nonparametric investigates the effects of China joining the WTO, the so-called China shock. Specifically, the authors examine what would have happened to China's welfare if China's trade costs had stayed constant at their 1995 levels. They consider $n$ countries and $T$ time periods. The exercise I am considering is assessing the sensitivity of counterfactual predictions to measurement error in bilateral trade flows.
The counterfactual objects of interest is the change in China's welfare, defined as the percentage change in income that the representative agent in China would be indifferent about accepting instead of the counterfactual change where China's trade costs are fixed at their 1995 levels. The details of the model can be found in Appendix (ref).\footnote{In adao2017nonparametric, the authors consider two demand systems: standard CES and “Mixed CES”. I focus on the standard CES specification.} The key insight is that we can express the proportional change in China's welfare in period $t$, denoted $W_{\mathrm{China},t}^{\mathrm{cf,prop}}$, as a function of the full set of bilateral trade flows across periods, $\left\{ F_{ij,t}\right\} $, and the trade elasticity, $\varepsilon$. Implementing this mapping requires multiple years of data, since answering the counterfactual question necessitates recovering the proportional change in China’s trade costs between 1995 and year $t$. This change is inferred from bilateral trade flows observed in those two years. Hence, we can write
for $t=1,...,T$ and known functions $g_{t}:\mathbb{R}_{+}^{Tn\left(n-1\right)}\times\mathbb{R}_{++}\rightarrow\mathbb{R}$. Then, conditional on a prior distribution for the true bilateral flows $\left\{ F_{ij,t}\right\} $ and a measurement error model, we can quantify uncertainty for $\left\{ W_{\mathrm{China},t}^{\mathrm{cf,prop}}\right\} $.
The default approach from Section (ref) can be applied. For the empirical Bayes step, the calibration of $\vartheta$, we can use the mirror trade data setting from Section (ref). Since there are no zero flows in this application, the estimated posterior of interest is \[ F_{ij,t}|\tilde{F}_{ij,t}\sim\exp\left\{ \mathcal{N}\left(\frac{\mathring{s}_{ij}^{2}}{\mathring{s}_{ij}^{2}+\mathring{\varsigma}_{ij}^{2}}\log\left(\tilde{F}_{ij,t}\right)+\frac{\mathring{\varsigma}_{ij}^{2}}{\mathring{s}_{ij}^{2}+\mathring{\varsigma}_{ij}^{2}}\tilde{\mu}_{ij,t},\left(\frac{1}{\mathring{s}_{ij}^{2}}+\frac{1}{\mathring{\varsigma}_{ij}^{2}}\right)^{-1}\right)\right\} , \] where $\left\{ \mathring{s}_{ij}^{2}\right\} $, $\left\{ \mathring{\varsigma}_{ij}^{2}\right\} $, $\left\{ \tilde{F}_{ij,t}\right\} $ and $\left\{ \tilde{\mu}_{ij,t}\right\} $ are all defined in Appendix (ref).
Having obtained a posterior distribution for the true trade flows given the noisy trade flows, we can now quantify uncertainty about the counterfactual predictions of interest. In Figure (ref), I reproduce Figure 3 of adao2017nonparametric, which plots the percentage change in China's welfare as a result of the China shock for each year in the period 1996-2011, and include two $95\%$ intervals.
The first region only considers estimation error and hence assumes the data are perfectly measured. It is constructed using code provided by the authors, and samples from the normal distribution with mean and variance equal to the GMM estimator for the trade elasticity $\varepsilon$ and its sampling variance, respectively. The resulting intervals are small for the period before the year 2000, and then slowly become wider. These are the intervals reported in adao2017nonparametric.
The second region considers only measurement error and no estimation error in $\varepsilon$. The resulting intervals are considerably wider than the intervals based on estimation error, especially in the first few years. In Appendix (ref) I provide additional discussion and analyses.
The empirical application in allen2022welfare aims to estimate the returns on investment for all highway segments of the US Interstate Highway network. The authors do so by introducing an economic geography model and calculating what happens to welfare after a $1\%$ improvement to all highway links. Combining these counterfactual welfare changes with how many lane-miles must be added in order to achieve the $1\%$ improvement, they find the highway segments with the greatest return on investment.
This exercise only requires data on incomes and traffic flows of the $n$ locations and knowledge of four structural model parameters. The details of the model can be found in Appendix (ref), but the key relation is the one that maps the average annual daily traffic (AADT) flows $\left\{ F_{ij}\right\} $ to the counterfactual return on investments $\left\{ R_{k\ell}^{\mathrm{cf}}\right\} $, which is \[ R_{k\ell}^{\mathrm{cf}}=g_{k\ell}\left(\left\{ F_{ij}\right\} ,\theta\right) \] for known functions $g_{k\ell}:\mathbb{R}_{+}^{n\left(n-1\right)}\times\Theta\rightarrow\mathbb{R}$ for $k,\ell=1,...,n$ locations in the US Interstate Highway network.
For this application we can again apply the default approach from Section (ref). For the empirical Bayes step we can use the baseline case from Section (ref). There are no zeros so we only have to provide an estimate of the measurement error variance $\varsigma^{2}$. musunuru2019applications estimates that the measurement error variance of the logarithm of the average annual daily traffic (AADT) flows, which is exactly the data that allen2022welfare uses, is between 0.05 and 0.20. To obtain a lower bound on uncertainty, I will use a uniform measurement error variance of $0.05$.
With $\tilde{\varsigma}^{2}=0.05$, I use Equation (ref) to find a prior variance of $\tilde{s}^{2}=0.101$. This results in the following estimated posterior distribution for the true traffic flow between country $i$ and $j$, $F_{ij}$, given its noisy version $\tilde{F}_{ij}$, for $i,j=1,...,n$: \[ F_{ij}|\tilde{F}_{ij}\sim\mathrm{exp}\left\{ \mathcal{N}\left(0.669\cdot\log\tilde{F}_{ij}+0.331\cdot\tilde{\mu}_{ij},0.033\right)\right\} , \] where $\tilde{\mu}_{ij}$ is defined in Equation (ref).
The counterfactual question of interest is which links have the highest return on investment, and the authors of allen2022welfare report the top ten links. For exposition, I will focus my analysis on the three best performing links. Table (ref) shows the $95\%$ intervals for the top three links based on Algorithm (ref).
From a policy perspective it is of interest whether the ranking between these links can change due to measurement error. Therefore, Table (ref) shows the $95\%$ intervals for the difference between link 1 and link 2, and the difference between link 2 and link 3.\footnote{This simple exercise is intended purely for exposition. For a more formal treatment of inference on ranks, see mogstad2024inference.} It follows that the rankings are generally robust against measurement error. Additional discussion and analyses can be found in Appendices (ref) and (ref).
This paper develops an econometric framework for quantifying the impact of measurement error in a broad class of quantitative trade and spatial models. Unlike standard econometric models of measurement error, the counterfactual estimand in these models depends directly on the realized data rather than on the underlying distribution. I adopt an empirical Bayes approach to characterize uncertainty in counterfactual predictions and propose a default specification that can be easily implemented across a range of applications. Applying the framework to the settings in adao2017nonparametric and allen2022welfare, I find substantial uncertainty surrounding key economic outcomes. These results underscore the need to account for measurement error when using quantitative models to guide policy decisions.