Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
642,264 characters · 14 sections · 56 citation commands
A Framework for Eliciting, Incorporating, and Disciplining Identification Beliefs in Linear Models
\affil[1]{ Department of Economics, University of Oxford} \affil[2]{ Federal Reserve Bank of Chicago & NBER}
\thispagestyle{empty}
\setcounter{page}{1}
\epigraph{“Belief is so important! A hundred contradictions might be true.”}{--- Blaise Pascal, Pens\'{e}es}
To identify causal effects from observational data, an applied researcher must augment the data with her beliefs. The exclusion restriction in an instrumental variables (IV) regression, for example, represents the belief that the instrument has no direct effect on the outcome of interest. Even when this belief cannot be tested directly, applied researchers know how to think about it and how to debate it. In practice, however, not all beliefs are treated equally. In addition to “formal beliefs” such as the IV exclusion restriction -- beliefs that are directly imposed to obtain identification -- researchers often state a number of “informal beliefs.” While not directly imposed on the problem, informal beliefs play an important role in interpreting results and reconciling conflicting estimates. Papers that report IV estimates, for example, almost invariably state the authors' belief about the sign of the correlation between the endogenous treatment and the error term but do not exploit this information in estimation.\footnote{Referring to more than 60 papers published in the top three empirical journals between 2002 and 2005, Moon2009 note that “in almost all of the papers the authors explicitly stated their beliefs about the sign of the correlation between the endogenous regressor and the error term; yet none of the authors exploited the resulting inequality moment condition in their estimation.”} Another common informal belief concerns the extent of measurement error. When researchers observe an ordinary least squares (OLS) estimate that is substantially smaller than, but has the same sign as its IV counterpart, classical measurement error, with its attendant “least squares attenuation bias,” is often suggested as the likely cause.
Relegating informal beliefs to second-class status is both wasteful of information and dangerous; beliefs along different dimensions of the problem are mutually constrained by each other, the model, and the data. By failing to explicitly incorporate all relevant information, applied researchers both leave money on the table and, more importantly, risk reasoning to a contradiction by expressing mutually incompatible beliefs. Although this point is general, we illustrate its implications here in the context of a linear model
where $T^*$ is a potentially endogenous treatment, $y$ is an outcome of interest, and $\mathbf{x}$ is a vector of exogenous controls. Our goal is to estimate the causal effect of $T^*$ on $y$, namely $\beta$, but we observe only $T$, a noisy measure of $T^*$ polluted by measurement error $\widetilde{w}$. While we are fortunate to have an instrument $z$ at our disposal, it may not satisfy the exclusion restriction: $z$ is potentially correlated with $u$. This scenario is typical in applied work: endogeneity is the rule rather than the exception, the treatments of greatest interest are often the hardest to measure, and the validity of a proposed instrument is almost always debatable.
We focus on two cases that are common in applied work. In the first $T^*$ has no support restrictions and is subject to classical measurement error. In the second $T^*$ is binary and thus any errors in measurement must be non-classical.\footnote{ If $T^* = 1$, the only way it can be mis-measured is downwards: $T=0$. If $T^* = 0$ the only way it can be mis-measured is upwards: $T=1$. Hence $\widetilde{w}$ must be negatively correlated with $T^*$.} To accommodate both cases within a single framework, we derive our results under the assumption that $\widetilde{w}$ is non-differential. This permits correlation between $\widetilde{w}$ and $T^*$ but imposes the restriction that $\widetilde{w}$ is uncorrelated with all other random variables in the system conditional on $T^*$. We begin by deriving the sharp identified set relating treatment endogeneity, instrument invalidity, and non-differential measurement error when $T^*$ has unrestricted support. To the best of our knowledge, this result is new to the literature. Turning our attention to the binary $T^*$ case, we then show that adding support restrictions provides additional identifying information via cross-parameter restrictions. In both cases, however, the data alone provide no restrictions on $\beta$. As such, the addition of researcher beliefs is unavoidable. Using our characterization of the identified set, we propose a framework for Bayesian inference for the treatment effect of interest that combines the data with researcher beliefs in a coherent and transparent way. As we show in our empirical examples, this framework not only allows researchers to incorporate relevant problem-specific beliefs, but helps them to refine and discipline them by revealing any inconsistencies that may be present.
Whenever one imposes information beyond what is contained in the data, it is crucial to make clear how this information affects the ultimate result. Accordingly, we decompose our problem into a vector of partially-identified structural parameters $\boldsymbol{\theta}$, and a vector of point-identified reduced form parameters $\boldsymbol{\varphi}$. The vector $\boldsymbol{\theta}$ contains the parameters that govern instrument invalidity, regressor endogeneity and measurement error, while $\boldsymbol{\varphi}$ contains observable moments obtained from reduced form regressions of $(y,T,z)$ on $\mathbf{x}$. This decomposition is structured so that the data are only informative about $\boldsymbol{\theta}$ through $\boldsymbol{\varphi}$, revealing precisely how any identification beliefs we may choose to impose enter the problem.\footnote{Such a decomposition is called a transparent parameterization in the statistics literature. See, for example GustafsonBookx.} In particular, the data rule out certain values of $\boldsymbol{\varphi}$, while our beliefs place restrictions on the conditional identified set $\Theta(\boldsymbol{\varphi})$ for $\boldsymbol{\theta}$. A prior over the conditional identified set $\Theta(\boldsymbol{\varphi})$ will never be updated by any amount of data. For this reason, prior elicitation for $\boldsymbol{\theta}$ is particularly crucial. Our approach to elicitation for $\boldsymbol{\theta}$ has two components. First, we parameterize measurement error, regressor endogeneity, and instrument invalidity in terms of intuitive, empirically meaningful parameters: correlations and what is in essence a signal-to-noise ratio. Second, because it can be challenging for researchers to articulate fully informative prior information, we consider only relatively weak prior beliefs in the form of sign and interval restrictions on the components of $\boldsymbol{\theta}$. These are fairly easy to elicit in practice and can be surprisingly informative about the causal effect of interest. We present two complementary approaches to Bayesian inference for the structural parameters: inference for the identified set $\Theta$, and inference for the partially identified parameter $\boldsymbol{\theta}$ under a conditionally uniform reference prior. We compare and contrast these approaches below.
While measurement error, treatment endogeneity, and invalid instruments have all generated voluminous literatures, to the best of our knowledge this is the first paper to carry out a partial identification exercise in which all three problems can be present simultaneously. Our main point is simple but has important implications for applied work that have been largely overlooked; measurement error, treatment endogeneity, and instrument invalidity are mutually constrained by each other and the data in a manner that can only be made apparent by characterizing the full identified set for the model. Because the dimension of this set is strictly smaller than the number of variables used to describe it, the constraints of the model could easily contradict prior researcher beliefs. Given the shape of the identified set, the belief that $z$ is a valid instrument, for example, could imply an implausible amount of measurement error or a selection effect with the opposite of the expected sign. In this way our framework provides a means of reconciling and refining beliefs that would not be possible based on introspection alone. We are by no means the first to recognize the importance of requiring that beliefs be compatible. KahnemanTversky, for example, make a closely related point in their discussion of heuristic decision-making under uncertainty. Even if specific probabilistic assessments appear coherent on their own,
Our purpose here is to take up the challenge laid down by KahnemanTversky and provide just such a formal procedure for assessing the compatibility of researcher beliefs over treatment endogeneity, measurement error, and instrument invalidity in linear models. Although the intuition behind our procedure is straightforward, the details are more involved. For this reason we provide free and open-source software in R to make it easy for applied researchers to implement the methods described in this paper.\footnote{See \url{https://github.com/fditraglia/ivdoctr}.}
This paper contributes to a small but growing literature on the Bayesian analysis of partially-identified models, including Poirier1998, Richardsonetal, Moon2012, Hahnetal, and GustafsonBookx. Some recent contributions to the literature on structural vector autoregression models BaumeisterHamilton, Arias, AmirAhmadi also explore related ideas. Because we discuss, as part of our exercise, Bayesian inferences for the identified set, our work relates to Kitagawa, KlineTamer, and TimBayes who give sufficient conditions under which such inferences have a valid frequentist interpretation.
Our results relate to the classical literature on errors in variables in linear models, for example KlepperLeamer, Leamer, and bekker1987. The main distinction between our paper and this literature is threefold. First, our regressor of interest $T^*$ is endogenous; second, the measurement error $\widetilde{w}$ that generates our observed regressor $T$ may be non-classical; third we consider settings in which a (potentially imperfect) instrumental variable is available. While the proxy variable setting considered in KraskerPratt and Bollinger2003 can be interpreted as a non-classical measurement error problem, these papers likewise consider only exogenous regressors and do not rely on an instrumental variable. Our results also relate to a large literature on estimating the effect of mis-measured binary regressors without relying on instrumental variables. An early contribution is Bollinger who provides partial identification bounds for the effect of an exogenous, binary regressor subject to non-differential mis-classification. HasseltBollinger derive additional bounds for the same model. BollingerHasseltWP propose a Bayesian inference procedure based on these bounds and consider an extension that addresses potential endogeneity in the true, unobserved regressor by placing a prior on its covariance with the error term. In contrast, kreider2007, kreider2012, and gundersen2012 derive partial identification bounds for the effect of a binary regressor subject to arbitrary mis-classification error when the outcome of interest is also binary. The latter two papers allow for endogeneity in the true, unobserved regressor.
Because we consider a situation in which an instrumental variable is available, our setting is more closely related to that considered by KRS, BBS, FL, Lewbel, Mahajan and Hu2008. The key lesson from these papers is that the two-stage least squares (TSLS) estimator is inconsistent even if the instrument is valid. When the treatment is exogenous, however, it is possible to construct a non-linear method of moments estimator that recovers the treatment effect using a discrete instrumental variable. Unlike these papers, we consider a setting in which the binary treatment of interest may be endogenous. As shown in DiTragliaGarciaJimeno_b the usual instrumental variable assumption is insufficient to identify the effect of an endogenous, mis-measured, binary treatment. While that paper provides a point identification result under a stronger instrument exclusion restriction, we do not rely on it here. Instead we allow for an invalid instrument and derive partial identification bounds.
Two papers that similarly consider partial identification under instrument invalidity are Conley2012 and Nevo2012. Like us, Conley2012 adopt a Bayesian approach that allows for a violation of the IV exclusion restriction, but they do not explore the relationship between treatment endogeneity and instrument invalidity. In contrast, Nevo2012 derive bounds for a causal effect in the setting where an endogenous regressor is “more endogenous” than the variable used to instrument it is invalid. Our framework encompasses the settings considered in these two papers, but is strictly more general in that we allow for measurement error simultaneously with treatment endogeneity and instrument invalidity. More importantly, the central message of our paper is that it can be misleading to impose beliefs on only one dimension of a partially identified problem unless one has a way of ensuring their mutual consistency with all other relevant researcher beliefs. For example, although a single valid instrument solves both the problem of classical measurement error and treatment endogeneity, it is insufficient to carry out a partial identification exercise that merely relaxes the exclusion restriction, as in Conley2012. Values for the correlation between $z$ and $u$ that seem plausible when viewed in isolation could easily imply implausible amounts of measurement error or treatment endogeneity.
The remainder of this paper is organized as follows. Section (ref) derives the sharp identified set when $T^*$ has unrestricted support. Section (ref) considers the case in which $T^*$ is binary, deriving additional cross-parameter restrictions that apply in this setting. Section (ref) details our two approaches to Bayesian inference, including details of prior elicitation, using the results of Sections (ref) and (ref). Section (ref) presents a number of substantive empirical examples illustrating our procedure in both the classical measurement error and binary $T^*$ cases, and Section (ref) concludes. Proofs, auxiliary results, and additional computational details appear in an online appendix.
In this section we derive the joint restrictions relating measurement error, regressor endogeneity, and instrument invalidity given the observed data. We then use these restrictions to show how the identified set for $\beta$ depends on researcher beliefs over the three dimensions. Our approach is as follows. First, we use the assumption of non-differential measurement error to re-write (ref) in terms of a classical measurement error component $w$ and a parameter $\psi$ that governs the “non-classical” part of measurement error, an approach similar to that followed by Bollinger2003 in a proxy-variable setting.
Second, we relate the structural model from (ref)--(ref) to a system of reduced form regressions of $(y,T,z)$ on $\mathbf{x}$. The restrictions that we use in our partial identification exercise below arise from the mapping between structural and reduced form covariance matrices, along with the assumption of non-differential measurement error. Third, we re-parameterize our problem to “absorb” the non-classical measurement error parameter $\psi$. This allows us to proceed as though the measurement error were classical, and adjust for $\psi$ in a second step, greatly simplifying the calculations. The bounds we derive in this section are sharp provided that $T^*$ has full support. When the support of $T^*$ is restricted, however, it may be possible to tighten them, a possibility that we explore for a binary $T^*$ in (ref) below.
We begin by stating the basic assumptions that will be used throughout the paper.
The only substantive restrictions in (ref) are (i) and (v): (i) assumes that the control regressors $\mathbf{x}$ are exogenous, while (v) assumes that the mis-measured regressor $T$ is positively correlated with the true, unobserved regressor $T^*$. (ref) (ii) can be taken as the definition of the error term $v$ from (ref). It equals the residual from a projection of the unobserved regressor of interest $T^*$ on the instrument $z$ and exogenous control regressors $\mathbf{x}$. (ref) (iii) is the standard instrumental variables relevance condition, but stated for the unobserved true regressor $T^*$ rather than the observed, mis-measured regressor $T$. Although $T^*$ is unobserved, (ref) (iii) is testable under our other assumptions.\footnote{See (ref) and the discussion immediately following it for details.} Throughout this paper we will abstract from weak instrument considerations.
The main additional assumption that we rely on below concerns the nature of the measurement error $\widetilde{w}$ from (ref).
(ref) requires that any correlation between $\widetilde{w}$ and $(u,z,\mathbf{x})$ arises solely from correlation between $T^*$ and $(u,z,\mathbf{x})$. In other words we assume that $T$ contains no additional information about $(u,z,\mathbf{x})$ beyond that contained in $T^*$. Non-differential measurement error is the natural generalization of classical measurement error to settings where $T$ and $T^*$ have restricted support. As such, it is widely used in the literature on mis-classified discrete variables Lewbel,Mahajan,FL,Hu2008,DiTragliaGarciaJimeno_b. When $\psi = 0$, (ref) reduces to the classical case. When $\psi \neq 0$ it generalizes classical measurement error by allowing $\widetilde{w}$ to be correlated with $T^*$. This extra generality is necessary if we wish to consider a binary $T^*$ because $\widetilde{w}$ must be correlated with $T^*$ in this case: if $T^*=1$ then $\widetilde{w}$ must be $0$ or $-1$; if $T^*=0$ then $\widetilde{w}$ must be $0$ or $1$. (ref) places no restriction on the conditional distribution of $T$ given $T^*$ and hence no restriction on $\psi$; it merely imposes that $T$ is exogenous after projecting out $T^*$. This is indeed a restriction, but a strictly weaker one than classical measurement error.
Before proceeding, we require some additional notation. First let
where $\psi$ is as defined in (ref). Using (ref), we can re-write (ref) as
where $(1 + \psi) > 0$ by (ref) (v), to ensure that $T$ is positively correlated with $T^*$. Both (ref) and (ref) are completely without loss of generality: (ref) can be viewed as the definition of $\widetilde{w}$ and (ref) as the corresponding definition of $w$. Because $w$ is defined as the residual from a projection of $\widetilde{w}$ onto $T^*$ and a constant, it has zero mean and is uncorrelated with $T^*$ by construction, making (ref) more convenient to work with than (ref). In contrast, $\widetilde{w}$ may have a non-zero mean and be correlated with $T^*$. Although $T$ and $T^*$ are positively correlated by (ref) (v), note that the correlation between $T^*$ and $\widetilde{w}$ may be positive or negative as $\psi \in (-1, +\infty)$.
At the heart of our partial identification exercise is the relationship between reduced form and structural covariance matrices. Define the reduced form model as
where $(\varepsilon, \xi, \zeta)$ are projection errors with covariance matrix
Under (ref) $(y,T,z,\mathbf{x})$ are observed, so $\Phi \equiv (\boldsymbol{\varphi}_y , \boldsymbol{\varphi}_T, \boldsymbol{\varphi}_z)$ and $\Sigma$ are point identified. Throughout the paper, we will refer to $\Sigma$ as the reduced form covariance matrix. To avoid trivial but uninteresting cases, we assume throughout that $\Sigma$ is positive definite. Let $\Omega$ denote the covariance matrix of $(u,v,\zeta,w)$. We will refer to $\Omega$ as the structural covariance matrix.\footnote{Note that our convention treats $\zeta$ as both a structural and reduced form error.} $\Omega$ is unobserved because $T^*$ is unobserved and potentially endogenous. We assume that $\Omega$ is “well-behaved” in the following sense.
(ref) does not require that $\Omega$ be positive definite. This allows for the possibility that there is no measurement error, in which case $\mbox{Var}(w) = 0$. Note that we treat $w$ rather than $\widetilde{w}$ as the “structural” measurement error. The advantage of following this convention is that $w$, unlike $\widetilde{w}$, satisfies all of the assumptions of classical measurement error, as shown in the following lemma.
(ref) allows for the possibility that $z$ is an invalid instrument, $\sigma_{u\zeta} \neq 0$, and that $T^*$ is endogenous, $\sigma_{uv} \neq 0$. The zeros in $\Omega$ arise from (ref) (ii), which ensures that $v$ is uncorrelated with $\zeta$, and (ref), which ensures that $w$ has the properties of classical measurement error. We now turn our attention to the relationship between the reduced form covariance matrix $\Sigma$ and the structural covariance matrix $\Omega$. This relationship emerges as a corollary of the following lemma.
(ref) shows that the reduced form coefficients $\boldsymbol{\varphi}_T$ and $\boldsymbol{\varphi}_y$ are functions of the structural parameters $(\beta, \pi, \psi)$. While it may appear from this result that knowledge of $(\boldsymbol{\varphi}_y, \boldsymbol{\varphi}_T, \boldsymbol{\varphi}_z)$ provides additional identifying information, this is not the case. Given values for the reduced form regression coefficients $(\boldsymbol{\varphi}_y, \boldsymbol{\varphi}_T, \boldsymbol{\varphi}_z)$, we can construct values of the structural regression coefficients $\boldsymbol{\eta}$ and $\boldsymbol{\gamma}$ that are consistent with any desired values of the other structural paramters, namely
where (ref) (v) justifies division by $(1 + \psi)$: if $\mbox{Cov}(T,T^*) > 0$ then $\psi > -1$ as seen from (ref). More importantly, (ref) implies that $\Sigma$ is related to $\Omega$ according to
Expanding $\Sigma = \Gamma \Omega \Gamma'$, we obtain the following:
Equations (ref)--(ref) constitute the restrictions that we will use to carry out our partial identification exercise below. (ref) reveals that (ref) (iii), instrument relevance, is testable: $(1 + \psi)\pi = (s_{23}/s_{33})$ and $(1 + \psi)$ cannot equal zero by (ref) (v). As shown in the following lemma, however, Assumptions (ref)--(ref) and the relationship $\Sigma = \Gamma \Omega \Gamma'$ impose no restrictions on the parameter $\psi$ other than $\psi > -1$.
(ref) shows that, without further restrictions, the reduced form covariance matrix contains no information about $\psi$. Indeed an even stronger result holds: unless $T^*$ has support restrictions, a model with structural parameters $\theta$ is observationally equivalent to one with structural parameters $\theta'$.\footnote{See the proof of (ref) for details.} Intuitively, because $T^*$ is unobserved we are free to arbitrarily re-scale both sides of (ref) -- effectively “redefining” $T^*$ -- so long as we absorb this rescaling into the remaining parameters of the system. If $T^*$ has a restricted support, however, such an arbitrary rescaling is no longer possible. For example, if $T^*$ is binary, certain choices of scale can be ruled out by observing the distribution of $T$. In this case it is still true that $\Sigma$ on its own contains no information about $\psi$, but the binary nature of $T^*$ creates additional cross-parameter restrictions that can be used to bound $\psi$. Because binary treatments are common in applied work, we develop this special case in full detail in (ref). Analogous reasoning applies to the parameter $\tau$ from (ref). Without support restrictions on $T^*$ we can shift $\tau$ arbitrarily while fixing $\mathbb{E}[T]$, absorbing the difference into $\mathbb{E}[T^*]$ and the first-stage intercept.
Before proceeding to derive the joint restrictions between measurement error, regressor endogeneity, and instrument invalidity, we first re-write equations (ref)--(ref) in a form that simplifies both our mathematical derivations and, ultimately, the elicitation of researcher beliefs. To begin, we define a reduced form regression for the unobserved regressor $T^*$. Using logic analogous to that of (ref), we can write
Since $\zeta$ is uncorrelated with $v$ by (ref) (ii), it follows that
Equation (ref) shows that endogeneity in $T^*$ arises from two sources: invalidity of the instrument $z$, and correlation between the error terms $u$ and $v$. By representing regressor endogeneity in terms of $\sigma_{u\xi^*}$, (ref) allows us to eliminate $\sigma_{uv}$ from (ref)--(ref). Next we define the parameter $\kappa$ as
where the last equality follows by solving (ref) for $(\pi^2 s_{33} + \sigma_v^2)$. In the special case where $\mathbf{x}$ includes only a constant, $T^*$ is exogenous, and the measurement error is classical, $\kappa$ measures the degree of attenuation bias present in the OLS estimator. More generally, $\kappa$ measures the proportion of “signal” contained in the reduced form error $\xi^*$. If $\kappa = 1/2$, for example, this means that half of the variation in $\xi$ is generated by $\xi^*$, and the remainder is “noise” arising from $w$. Unlike $\sigma_w^2$, $\kappa$ has bounded support: $\kappa \in (0,1]$. When $\kappa = 1$, $\sigma_w^2=0$ so there is no measurement error; the limit as $\kappa$ approaches zero corresponds to taking $\sigma_w^2$ to its maximum possible value: $s_{22}$. Finally, define
The parameters defined in (ref) correspond to setting $\psi' = 0$ in (ref), which “absorbs” the non-classical component of measurement error, $\psi$, into the definitions of the remaining parameters. Note that if the measurement error $\widetilde{w}$ is in fact classical, then $\psi = 0$ so that $\widetilde{\beta} = \beta$, $\widetilde{\beta} = \pi$, and so on. Using (ref)--(ref), we can re-write (ref)--(ref) as
In essence, we have transformed a problem with non-classical measurement error into an equivalent problem with classical measurement error but different parameter values. In the transformed system, the extent of measurement error is controlled by $\widetilde{\kappa}$ and regressor endogeneity is controlled by $\widetilde{\sigma}_{u\xi^*}$. Instrument invalidity is controlled by the same parameter in both the original and transformed parameterizations: $\sigma_{u\zeta}$. While $\widetilde{\kappa}$ is scale-free, $\sigma_{u\zeta}$ and $\widetilde{\sigma}_{u\xi^*}$ are not. For this reason, when we derive the restrictions implied by (ref)--(ref) below we will express them in terms of correlations rather than covariances, namely
Note that
so that $\rho_{u\xi^*}$, unlike $\sigma_{u\xi^*}$, is unaffected by the re-parameterization in (ref)--(ref). In summary, we can proceed as though the measurement error were classical by working in terms of $(\rho_{u\zeta}, \rho_{u\xi^*}, \widetilde{\kappa})$. Any restrictions on $\psi$, for example in the case of a binary $T^*$, can be addressed in a second step. In the following section, we derive the joint restrictions between these parameters and the identified set for $\beta$.
A key point of this paper is that beliefs over measurement error, regressor endogeneity, and instrument invalidity are mutually constrained by each other and the data. The following result makes this intuition precise by expressing $\rho_{u\zeta}$ as an explicit function of $\rho_{u\xi^*}$ and $\widetilde{\kappa}$, given particular values of the reduced form correlations.
(ref) is the first ingredient in our characterization of the joint restrictions between measurement error, regressor endogeneity, and instrument invalidity. The second is a bound on $\widetilde{\kappa}$ that limits the possible extent of measurement error in the data.
Because it places a lower bound on $\widetilde{\kappa}$, namely $L$, (ref) places an upper bound on the extent of measurement error. The derivation of this bound relies on two simpler but weaker bounds. The first, $\widetilde{\kappa} > r_{12}^2$, corresponds to the familiar “reverse regression bound” under classical measurement error. The second, $\widetilde{\kappa} > r_{23}^2$, is in essence a reverse regression bound constructed from the IV first-stage. The bound $\widetilde{\kappa} > L$ is strictly tighter than both of these bounds, as it incorporates information from all three of the reduced form correlations: $r_{12}, r_{23}$, and $r_{13}$. (ref) does not, however, allow us to rule out the possibility that there is no measurement error: $\widetilde{\kappa} = 1$ always satisfies the bounds regardless of the values of the reduced form correlations.
Together, (ref) and (ref) provide joint restrictions on instrument invalidity, regressor endogeneity, and measurement error. In particular, the reduced form covariance matrix $\Sigma$ both bounds $\widetilde{\kappa}$ and gives $\rho_{u\zeta}$ as an explicit function of $\rho_{u\xi^*}$ and $\widetilde{\kappa}$. These restrictions in fact constitute the sharp identified set, as we now show.
The additional assumption $s_{23} \neq 0$ in (ref) is a reduced form version of the structural instrument relevance condition from (ref) (iii); it requires that $z$ is correlated with $T$ even after projecting out $\mathbf{x}$. Note that (ref) imposes no cross-restrictions between the parameters $\widetilde{\kappa}$, $\psi$, $\tau$, and $\rho_{u\xi^*}$. In contrast, $\rho_{u\zeta}$ is completely determined by $\widetilde{\kappa}$ and $\rho_{u\xi^*}$ by (ref). Moreover, $\psi$, $\tau$ and $\rho_{u\xi^*}$, unlike $\widetilde{\kappa}$, are completely unrestricted by observables. As shown in the following result, our assumptions also bound the instrument invalidity parameter $\rho_{u\zeta}$, despite placing no restriction on regressor endogeneity.
Because $L > r_{23}^2$, (ref) always rules out a range of values for $\rho_{u\zeta}$. Notice, however, that it never rules out $\rho_{u\zeta} = 0$. This is unsurprising given that it is known to be impossible to test for instrument validity in the model we consider here. Unfortunately, and also unsurprisingly, the model itself places no restrictions on the causal effect $\beta$.
The only way to learn about $\beta$ in this model is to impose beliefs. In our examples below we consider simple interval restrictions on $\widetilde{\kappa}$ and $\rho_{u\xi^*}$. (ref) in the appendix shows how interval restrictions on $\widetilde{\kappa}$ and $\rho_{u\xi^*}$ tighten the bounds for $\rho_{u\zeta}$ from (ref). (ref) shows that any restriction on $\rho_{u\xi^*}$ that rules out values arbitrarily close to -1 or 1 yields finite bounds for $\widetilde{\beta}$. In the case of classical measurement error, $\psi = 0$ and hence bounds for $\widetilde{\beta}$ are equivalent to bounds for $\beta$. In the general case, translating bounds for $\widetilde{\beta}$ into bounds for $\beta$ requires restrictions on $\psi$. When $T^*$ is binary, the data provide such restrictions. In the following section we derive these restrictions and show how to incorporate them into our partial identification exercise.
In many applied studies the regressor of interest is binary: $T^*,T \in \left\{ 0,1 \right\}$. In this case (ref) no longer applies: the data impose additional restrictions on $\psi$ through the support restriction on $T^*$. We now show how to extend our analysis from (ref) to incorporate the additional information available in the binary $T^*$ case. Similar reasoning can be applied when $T^*$ has an arbitrary discrete support set, although we do not pursue the general case here. To begin, we define some additional notation specific to the binary setting. First let $p^* \equiv \mathbb{P}(T^*=1)$ and $p \equiv \mathbb{P}(T=1)$. Next define the mis-classification error rates $\alpha_0$ and $\alpha_1$ as follows:
The parameter $\alpha_0$ equals the probability of an upwards mis-classification error, observing $T=1$ when $T^*=0$. In contrast, $\alpha_1$ equals the probability of a downwards mis-classification error, observing $T=0$ when $T^*=1$. Using this notation, we can express $\psi, \tau$ and $w$ as functions of $(\alpha_0, \alpha_1)$ as follows.
(ref) reveals two important features of the binary $T^*$ case. First, while $\psi$ could be positive or negative in the general case, it must be negative in the binary case. Second, while $\tau$ and $\psi$ are in general two free parameters, they are linked through their joint dependence on $\alpha_0$ in the binary case. Under (ref) (v), we have $\psi > -1$. By (ref) this is equivalent to $\alpha_0 + \alpha_1 < 1$ when $T^*$ is binary. The following Lemma exploits this fact to relate $p^*$ to $p$ and to yield a simple expression for $\sigma_w^2$ in terms of $(\alpha_0, \alpha_1)$ and $p$.
We now have two equations for $\sigma_w^2$ in the binary $T^*$ case: (ref), and (ref) (ii). Equating these yields the following cross-restriction between $\psi$ and $\widetilde{\kappa}$.
The intuition behind (ref) is as follows. In the binary $T^*$ case, both $\widetilde{\kappa}$ and $\psi$ are functions of the mis-classification probabilities $\alpha_0$ and $\alpha_1$. By definition these must lie between zero and one, and by (ref) (v) they also satisfy $\alpha_0 + \alpha_1 < 1$. This region is depicted in (ref). Since $\sigma_w^2 = s_{22}(1 - \widetilde{\kappa})$ by (ref), choosing a value for $\widetilde{\kappa}$ is equivalent to choosing a value of $\sigma_w^2$. Hence, solving the expression from (ref) (ii), the choice of $\widetilde{\kappa}$ determines $\alpha_1$ as a function of $\alpha_0$. The figure depicts three such functions, corresponding to three different choices of $\widetilde{\kappa}$: $L < \widetilde{\kappa}_1 < \widetilde{\kappa}_2$. Since $L < \widetilde{\kappa}$ by (ref), the first of these choices gives the outer envelope of this family of functions. The bounds for $\psi$ are determined by first pinning down a single function from this family by choosing a feasible value of $\widetilde{\kappa}$, and then finding all values of $C$ such that $\alpha_0 + \alpha_1 = C$ intersects this function. The minimum value of $(\alpha_0 + \alpha_1)$ always occurs at a corner. In the figure we set $p > 1/2$ so that the minimum occurs at $s_{22}(1 - \widetilde{\kappa})/p$. The maximum, indicated by the filled circles in the figure, can either be interior (red) or occur at a corner (blue). A corner maximum occurs when $\widetilde{\kappa}$ is sufficiently large, or equivalently $\sigma_w^2$ is sufficiently small. Finally, (ref) converts bounds for $(\alpha_0 + \alpha_1)$ into bounds for $\psi$.
In some cases, additional a priori information may be available to further restrict $\alpha_0$ and $\alpha_1$ and hence $(\widetilde{\kappa}, \psi)$. For example, under one-sided mis-classification, either $\alpha_0$ or $\alpha_1$ is known to be zero. Another such case is that of symmetric mis-classification, in which $\alpha_0 = \alpha_1$. A third example concerns settings in which auxiliary data suggest that $p^* \approx p$. This corresponds to the restriction $\alpha_1 \approx \alpha_0(1 - p)/p$. Each of these three special cases yields a linear equality restriction of the form $M_0 \alpha_0 + M_1 \alpha_1 = 0$ and reduces the number of unknown parameters by one. Geometrically this takes the form of a line with non-negative slope passing through the origin of (ref), meaning that $\psi$ is an explicit function of $\widetilde{\kappa}$. In the case of symmetric mis-classification, for example, $\psi$ is determined by the intersection of the 45-degree line and the curve corresponding to a given choice of $\widetilde{\kappa}$.
Without support restrictions, we know from (ref) that the data are uninformative about $\psi$. (ref) shows that when the support of $T^*$ is restricted to $\left\{ 0,1 \right\}$ this is no longer the case: the observables restrict $\psi$, and $\widetilde{\kappa}$ and $\psi$ are mutually constrained. (ref) in the Appendix shows how to use these restrictions to bound $\beta$. To summarize, the logic of (ref) shows that $\widetilde{\beta}$ is bounded so long as $\rho_{u\xi^*}$ is restricted a priori to lie in a strict subset of $(-1,1)$. (ref) combines this observation with (ref) to yield bounds for $\beta$ via (ref).
As we show in our empirical example from (ref) below, the restrictions imposed by (ref), in concert with (ref) and (ref), can be very informative in practice. Moreover, they allow us to treat the continuous and binary $T^*$ cases within a common, regression-based framework. However, these restrictions do not necessarily constitute the sharp identified set when $T^*$ is binary. For example, knowledge of the conditional distribution of $T|\mathbf{x}$ could in principle provide further restrictions on $(\alpha_0, \alpha_1)$. Exploiting this information, however, would require modeling objects over which applied researchers remain agnostic when reporting OLS and IV regressions, even with a binary $T^*$. Accordingly we do not purse this possibility further here.\footnote{For related results, see DiTragliaGarciaJimeno_b who derive the sharp identified set for a mis-classified, binary endogenous regressor given a valid instrument with discrete support, in an additively separable model with arbitrary dependence on exogenous covariates.}
We now describe how to use our results from above to carry out Bayesian inference. We present two approaches: inference for the identified set $\Theta$ and inference for the partially identified parameter $\boldsymbol{\theta}$. We focus throughout on two cases that are common in applications: first a regressor $T^*$ without support restrictions that is subject to classical measurement error, and second a binary $T^*$ as examined in (ref) above. Sections (ref) and (ref) consider classical measurement error, i.e.\ $\psi = 0$, in which case $\widetilde{\beta} = \beta$, $\widetilde{\kappa} = \kappa$, etc.\ Section (ref) explains the differences that arise when $T^*$ is binary. Online Appendix (ref) provides some discussion of the relationship between Bayesian and Frequentist inference in partially identified models.
Our approach relies on the principle that the choice of parameterization should make clear how any prior beliefs that cannot be falsified by data affect the ultimate result. For this reason, our derivations from above relate the identified set $\Theta$ for the structural parameters $\boldsymbol{\theta}$ to the reduced form parameters $\boldsymbol{\varphi} \equiv \left( \Sigma, \boldsymbol{\varphi}_y, \boldsymbol{\varphi}_T, \boldsymbol{\varphi}_z \right)$, i.e.\ $\Theta(\boldsymbol{\varphi})$, such that any inferences we draw about $\boldsymbol{\theta}$ depend on the data only through $\boldsymbol{\varphi}$.\footnote{This is called a transparent parameterization in the statistics literature: see, e.g., GustafsonBookx.} Because $\boldsymbol{\varphi}$ is point-identified, inference for this parameter vector is standard. We begin by assuming that the researcher has computed a posterior for $\boldsymbol{\varphi}$. Section (ref) discusses how to obtain one.
We elicit researcher beliefs in the form of sign and interval restrictions, $\mathcal{R}$, over regressor endogeneity, instrument invalidity, and measurement error. Intersecting $\Theta(\boldsymbol{\varphi})$ with $\mathcal{R}$ adds relatively weak prior information to restrict the identified set in a transparent manner. To simplify the elicitation of $\mathcal{R}$, our results in (ref) are expressed in terms of scale-free parameters. The regressor endogeneity parameter $\rho_{u\xi^*}$ and the instrument invalidity parameter $\rho_{u\zeta}$ are correlations, and have the same meaning regardless of whether the measurement error is classical or non-classical. In practice, a researcher might state a sign restriction for one or both of these quantities, along with an upper bound that is thought to represent an implausibly large extent of correlation. The appropriate way to elicit information about measurement error depends on the nature of that error. In the classical measurement error case $\psi = 0$ and hence $\kappa = \widetilde{\kappa}$. In this case, one could elicit interval restrictions over the scale-free variance ratio $\kappa$. Because $\kappa$ is defined net of covariates $\mathbf{x}$, it may be easier in some settings to instead elicit $\lambda \equiv \mbox{Var}(T^*)/\mbox{Var}(T)$ and transform this to $\kappa$ via $\kappa = (\lambda - R^2_{T.\mathbf{x}})/(1 - R^2_{T.\mathbf{x}})$ where $R^2_{T.\mathbf{x}}$ is the R-squared from a regression of $T$ on $\mathbf{x}$. In the binary $T^*$ case, neither $\psi$ nor $\widetilde{\kappa}$ is a natural parameter over which to elicit beliefs, but both are completely determined by $\alpha_0$ and $\alpha_1$. It is over these mis-classification probabilities, also scale-free, that researchers would most likely be able to state beliefs.
We first consider Bayesian posterior inference for the identified set for $\boldsymbol{\theta}$ rather than the structural parameter vector itself. If $\boldsymbol{\varphi}^{(j)}$ is a draw from the posterior for $\boldsymbol{\varphi}$, then $\Theta(\boldsymbol{\varphi}^{(j)})\cap \mathcal{R}$ is a draw from the posterior distribution for the identified set for $\boldsymbol{\theta}$ under researcher beliefs $\mathcal{R}$. By collecting a large number of these draws, one can summarize the posterior in a variety of different ways. First, one can construct a credible interval for the identified set of a particular structural parameter, such as $\rho_{u\zeta}$ or $\beta$, under a set of a priori restrictions $\mathcal{R}$. If $\mathcal{R}$ restricts $\rho_{u\xi}^*$ to a proper subset of $(-1,1)$, then (ref) yields two sided bounds for the instrument invalidity parameter $\rho_{u\zeta}$, while (ref) yields two-sided bounds for the causal effect $\beta$. Suppose we wish to form a 90% credible interval for the identified set $\mathscr{B}$ for $\beta$. To construct this interval, start with the conditional identified set $\mathscr{B}(\bar{\boldsymbol{\varphi}})$ evaluated at the posterior mean $\bar{\boldsymbol{\varphi}}$ and expand this interval outwards symmetrically until the resulting interval contains 90% of the identified sets. As we show in our empirical examples below, such intervals for $\beta$ can in some cases be surprisingly informative, despite relaxing the requirement that $z$ is a valid instrument.
Second, one can use the posterior to quantify the extent to which a particular set of a priori researcher beliefs $\mathcal{R}$ accords with the data by calculating the posterior probability that the intersection of $\Theta(\boldsymbol{\varphi})$ with $\mathcal{R}$ is empty. Consider, for example, a researcher who believes that selection is negative $(\rho_{u\xi^*}<0)$ and wishes to assess whether this is compatible with a belief that her instrument is valid $(\rho_{u\zeta} = 0)$. If we define $\mathcal{R}$ to be the intersection of these two restrictions, then calculating the fraction of sets $\Theta(\boldsymbol{\varphi}^{(j)}) \cap \mathcal{R}$ that are nonempty yields the posterior probability that we cannot rule out instrument validity under a particular assumption about the direction of selection. We abbreviate this as $\mathbb{P}(\text{Valid})$ in our empirical examples below. If $\mathbb{P}(\text{Valid})$ is small, the data strongly suggest that the assumed direction of selection is incompatible with instrument validity. More generally, consider any restriction $\mathcal{R}$. Calculating the fraction of sets $\Theta(\boldsymbol{\varphi}^{(j)}) \cap \mathcal{R}$ that are empty gives the posterior probability that $\mathcal{R}$ can be ruled out, a probability that we abbreviate as $\mathbb{P}(\varnothing)$ in our empirical examples below. If $\mathbb{P}(\varnothing)$ is small but nonzero, a researcher who feels confident in her a priori beliefs could elect to discard the draws $\boldsymbol{\varphi}^{(j)}$ for which $\Theta(\boldsymbol{\varphi}^{(j)})\cap \mathcal{R}$ is empty. If $\mathbb{P}(\varnothing)$ is large, this suggests that the beliefs encoded in $\mathcal{R}$ are suspect, given the data. When $\mathcal{R}$ restricts two or more dimensions of $(\rho_{u\zeta}, \rho_{u\xi^*}, \kappa)$, a large value of $\mathbb{P}(\varnothing)$ indicates that the corresponding researcher beliefs are mutually incompatible a posteriori. This exercise illustrates an important general point of our approach. By making explicit the relationship between measurement error, treatment endogeneity, and instrument invalidity, our method allows researchers to learn whether their beliefs over these different dimensions of the problem cohere.
Our second approach makes posterior probability statements about the partially identified parameter $\boldsymbol{\theta}$, by averaging both over reduced form draws $\boldsymbol{\varphi}^{(j)}$ and a conditional prior placed on $\Theta(\boldsymbol{\varphi}^{(j)})$. Carrying out inference for $\boldsymbol{\theta}$ rather than its identified set is attractive. For example, it allows one to compute the posterior probability that $\beta$ is positive. This, however, comes at a cost: the need to specify a conditional prior over the identified set. Because it may be difficult in practice to elicit a fully informative prior, following Moon2012 we recommend placing a uniform reference prior on $\Theta(\boldsymbol{\varphi}^{(j)})\cap \mathcal{R}$ (see Appendix (ref) for implementation details). Our use of this prior is intended to represent prior ignorance over $\Theta(\boldsymbol{\varphi}^{(j)})\cap\mathcal{R}$. Unavoidably, uniformity in one parameterization could imply a highly informative prior in some different parameterization. We emphasize, however, that the uniform serves here as a reference prior only. As such, one need not take it completely literally but could instead consider, for example, what kinds of deviations from uniformity would be necessary to support a particular belief about $\beta$.
A prior on the conditional identified set cannot be updated by the data. As such its influence on the posterior does not vanish as the sample size grows. For this reason, some caution is warranted when carrying out posterior inference for $\boldsymbol{\theta}$. A researcher who is concerned about this issue may wish to carry out a Bayesian robustness exercise over a class of priors supported on the conditional identified set. If this class includes all possible priors over $\Theta(\boldsymbol{\varphi}^{(j)})\cap \mathcal{R}$, the resulting bounds on posterior probabilities for $\boldsymbol{\theta}$ will coincide with our inferences for the identified set from (ref). While robust, such inferences are inherently conservative, as they summarize only the most extreme points of $\Theta(\boldsymbol{\varphi}^{(j)}) \cap \mathcal{R}$. Suppose for example that each draw $\Theta(\boldsymbol{\varphi}^{(j)}) \cap \mathcal{R}$ includes a single point that implies a negative value of $\beta$. Then, inference for the identified set $\mathscr{B}$ would produce no evidence against the claim that $\beta \leq 0$. In contrast, any reasonable prior over $\Theta(\boldsymbol{\varphi}^{(j)}) \cap \mathcal{R}$, such as our uniform reference prior, would give 100% posterior probability to $\{\beta > 0\}$.
We now summarize the modifications to our inference approaches from (ref) and (ref) that are required to treat the binary $T^*$ case from (ref). In this case, $\psi$ is in general non-zero and hence $\widetilde{\kappa}$ and $\widetilde{\beta}$ need not equal $\kappa$ and $\beta$. Note, however, that the meaning of $\rho_{u\zeta}$, along with that of $\rho_{u\xi^*}$, is unchanged in the binary $T^*$ case. Moreover, (ref) does not involve $\psi$, nor does (ref) impose cross-restrictions between $\rho_{u\zeta}$ and $\psi$. As such, to carry out inference for $\rho_{u\zeta}$ we can proceed exactly as we did in the classical measurement error case: all that changes is the interpretation of of $\widetilde{\kappa}$. This underscores a key advantage of working with a scale-free parameterization: the interpretations of $\rho_{u\zeta}$ and $\rho_{u\xi^*}$ do not depend on $\psi$. (ref) does, however, create a cross-restriction between $\widetilde{\kappa}$ and $\psi$. If $\widetilde{\beta}$ were our parameter of interest, we could ignore this fact and proceed as though the measurement error were classical. Because we are actually interested in $\beta = (1 + \psi)\widetilde{\beta}$, an extra step is required. To carry out inference for the identified set for $\beta$, we rely on (ref) to yield bounds for $\beta$ at any given reduced form draw $\boldsymbol{\varphi}^{(j)}$. To carry out inference for the partially identified parameter $\beta$, we first draw $\boldsymbol{\varphi}^{(j)}$ and then sample $(\widetilde{\kappa}^{(j)}, \rho_{u\xi^*}^{(j)}, \rho_{u\zeta}^{(j)})$ uniformly on the resulting conditional identified set, as described in (ref). We then draw $\psi^{(j)}$ uniformly from the interval $[\underline{\psi}(\widetilde{\kappa}^{(j)}), \overline{\psi}(\widetilde{\kappa}^{(j)})]$ defined in (ref). Given these draws, we construct the implied draw for $\beta^{(j)}$ using the derivations from (ref).
To implement the procedures from (ref) and (ref) the researcher must first obtain a posterior for the reduced form parameters. As we showed above in (ref), the reduced form regression slopes $(\boldsymbol{\varphi}_y,\boldsymbol{\varphi}_T,\boldsymbol{\varphi}_z)$ play no role in determining the identified set for $\boldsymbol{\theta}$. For this reason, we only require posterior draws for $\Sigma$. In our empirical examples below, we adopt the following simple approach. Given an iid sample of $n$ observations $\left( y_i, T_i, z_i, \mathbf{x}_i \right)$, let $\mathbf{y} = (y_1, \dots, y_n)'$ and define $\mathbf{T}$ and $\mathbf{z}$ analogously. Further define $X' = (\mathbf{x}_1', \dots, \mathbf{x}_n')$ and $Y = [
]$. We draw $\Sigma$ from an Inverse-Wishart$(\nu,S)$ distribution where \[ \nu = n - k + 3 + 1, \quad S = (Y - X\widehat{B})'(Y - X\widehat{B}), \quad \widehat{B} = (X'X)^{-1}X'Y \] and $k$ is the dimension of the exogenous covariate vector $\mathbf{x}_i$. Note that the mean of this distribution equals $S/(n - k)$, the sample covariance matrix of OLS residuals from the reduced form regressions given in \eqref{eq:reduced_yTz}. The Inverse-Wishart$(\nu,S)$ distribution is the marginal posterior for $\Sigma$ in the multivariate reduced form regression obtained by stacking (ref) under a Jeffreys prior and normal errors Zellner.
For simplicity, we draw the reduced form covariance matrix from an Inverse-Wishart posterior in both the classical measurement error and binary $T^*$ cases. Of course, the reduced form errors cannot be normal if any of the variables $(y,T,z)$ is discrete. Nonetheless, our Inverse-Wishart posterior for $\Sigma$ is still centered at $S/(n-k)$ and is approximately normal in large samples under mild conditions, as we discuss in Appendix (ref). Note that the bounds for $\psi$ from (ref) in the binary $T^*$ case involve $p$. To address this minor complication, we adopt an empirical Bayes approach, setting $p$ equal to the sample analogue $\widehat{p}$. Because this quantity is very precisely estimated, its effect on our inferences is negligible. An alternative to our Inverse-Wishart posterior for $\Sigma$ is the Bayesian Bootstrap approach followed by BollingerHasseltWP.
We now present three empirical examples illustrating how the framework described above can be applied in practice. The examples in Sections (ref) and (ref) involve a continuous treatment which we assume is subject to classical measurement error, i.e.\ $\psi = 0$, $\widetilde{\kappa} = \kappa$ and $\widetilde{\beta} = \beta$. In contrast, the example in Section (ref) involves a binary treatment, so that any measurement error that is present must be non-classical.
Acemoglu2001 study the effect of institutions on GDP per capita using a cross-section of 64 countries. Because institutional quality is endogenous, they use differences in the mortality rates of early western settlers across colonies as an instrumental variable. We consider their benchmark specification
which does not include covariates.\footnote{Additional results, available upon request, consider alternative specifications that include covariates. The results are essentially unchanged.} This yields an IV estimate of 0.94 with a standard error of 0.16 -- nearly twice as large as the corresponding OLS estimate of 0.52 with a standard error of 0.06. The authors attribute this disparity to classical measurement error:
Acemoglu2001 state two beliefs that are relevant for our partial identification exercise. First, their discussion implies there is likely a positive correlation between “true” institutions and the main equation error term $u$. This could arise from reverse causality -- wealthier societies can afford better institutions -- or omitted variables, such as legal origin or British culture, which are likely to be positively correlated with present-day institutional quality. We encode this belief using the prior restriction $0<\rho_{u\xi^*}<0.9$ below, ruling out only unreasonably large values of treatment endogeneity.\footnote{By (ref), the identified set for $\beta$ is $(-\infty,\infty)$ unless $\rho_{u\xi^*}$ is restricted. Here we impose the researchers' stated belief that $\rho_{u\xi^*}>0$ along with an extremely conservative upper bound for $\rho_{u\xi^*}$ of 0.9.} Second, in a footnote that uses an alternative measure of institutions as an instrument for the first, the authors argue that measurement error could be substantial.\footnote{Footnote \#19 of Acemoglu2001 states “We can ascertain, to some degree, whether the difference between OLS and 2SLS estimates could be due to measurement error by making use of an alternative measure of institutions \ldots This suggests that `measurement error' in the institutions variables \ldots is of the right order of magnitude to explain the difference between the OLS and 2SLS estimates.”} Taken at face value, the calculations from this footnote imply a point estimate of $\kappa = 0.6$ which would mean that 40 percent of the variation in measured institutions is noise.\footnote{Suppose $T_1$ and $T_2$ are two measures of institutions that are subject to classical measurement error: $T_1 = T^* + w_1$ and $T_2 = T^* + w_2$. Both $T_1$ and $T_2$ suffer from precisely the same degree of endogeneity, because they inherit this problem from $T^*$ alone under the assumption of classical measurement error. Thus, the OLS estimator based on $T_1$ converges to $\kappa(\beta+ \sigma_{T^*u}/\sigma_{T^*}^2)$ while the IV estimator that uses $T_2$ to instrument for $T_1$ converges to $\beta + \sigma_{T^*u}/\sigma_{T^*}^2$. The ratio identifies $\kappa$: $0.52/0.87 \approx 0.6$.} Below we consider two alternative ways of encoding this auxiliary information about $\kappa$.
Results for the Colonial Origins example appear in (ref). Estimates and bounds for $\beta$ indicate the percentage increase in GDP per capita that would result from a one point increase in the quality of institutions, as measured by average protection against expropriation risk. All other values in the table are unitless: they are either probabilities, correlations, or variance ratios. OLS and IV estimates and standard errors, along with an estimate of the lower bound $L$ for $\kappa$, appear in the first row of Panel (I). Panel (II) presents inferences for the identified set. The first column of Panel (II) gives the fraction of posterior draws for the reduced form parameters that yield an empty identified set, while the second column gives the fraction that are compatible with a valid instrument: $\rho_{u\zeta} = 0$. The third and fourth columns of Panel (II) present 90% posterior credible intervals for the identified sets for $\rho_{u\zeta}$ and $\beta$, constructed by symmetrically expanding around the conditional identified set evaluated at the posterior mean for $\Sigma$, as described in (ref). In contrast, panel (III) presents posterior medians and 90% highest posterior density intervals for $\rho_{u\zeta}$ and $\beta$, based on the uniform reference prior described in (ref).
We first consider an a priori restriction that $\kappa < 0.6$, placing a lower bound on the extent of measurement error. This restriction comes from personal communication with one of the authors of Acemoglu2001.\footnote{Based on footnote 19 of the paper, he expressed the belief that at least 40 percent of the measured variation in quality of institutions was likely to be noise.} Under this restriction, approximately 26 percent of the draws for the reduced form parameters yield an empty identified set, as shown in the first column of Panel (II). Intuitively, this means that there are covariance matrices $\Sigma$ that are close to the maximum likelihood estimate $\widehat{\Sigma}$ but which rule out the region $(\kappa, \rho_{u\xi^*}) \in (0,0.6]\times [0, 0.9]$. The problem is not the restriction on $\rho_{u\xi^*}$ but on $\kappa$: the data place no restrictions on the extent of treatment endogeneity although they do provide an upper bound on the extent of measurement error, as shown in (ref). Indeed, the proposed a priori upper bound of $0.6$ for $\kappa$ is only slightly larger than our point estimate of 0.54 for $L$, the lower bound defined in (ref). After accounting for uncertainty over $\Sigma$, we find that 26 percent of the posterior density for $L$ lies above 0.6. As such, our framework strongly suggests that the belief $\kappa < 0.6$ is incompatible with the data, and we cannot proceed further under this prior.
We now consider a second restriction that takes $0.6$ as a lower bound on $\kappa$, while continuing to impose $\rho_{u\xi^*} \in [0, 0.9]$. This restriction places an upper bound on the extent of measurement error, ruling out the most extreme possible values of $\kappa$. Results for this restriction appear in the third row of Table (ref). This restriction does not yield empty identified sets, as we see from the first column of Panel (II). It does however, strongly suggest that settler mortality is an invalid instrument: 70% of the posterior draws for the reduced form parameters exclude $\rho_{u\zeta}=0$ under the restriction $(\kappa, \rho_{u\xi^*}) \in (0.6,1]\times[0,0.9]$. Figure (ref) makes this point in a slightly different way, by depicting the identified set for $(\kappa, \rho_{u\xi^*}, \rho_{u\zeta})$, evaluated at the posterior mean for $\widehat{\Sigma}$, in the region where $\rho_{u\xi^*}$ is positive.\footnote{Note that under our Jeffreys prior the posterior mean equals the maximum likelihood estimator.} The gray region corresponds to $L < \kappa < 0.6$, the largest amount of measurement error consistent with $\widehat{\Sigma}$. We see from the figure that the plane $\rho_{u\zeta} = 0$ only intersects the identified set in the region where measurement error is extremely severe. Moreover, unless $\kappa = L$, $\rho_{u\zeta} = 0$ implies that $\rho_{u\xi^*}$ must be close to zero, in other words that institutions are approximately exogenous. This seems implausible. Indeed, under the restriction $(\kappa, \rho_{u\xi^*}) \in (0.6,1]\times[0,0.9]$, depicted in shades of red and blue in Figure (ref), the identified set resides exclusively below the plane $\rho_{u\zeta}=0$, suggesting that log settler mortality is negatively correlated with the unobservables.
Figure (ref) shows that one would need to place high a priori probability on implausible regions of the identified set to support the belief that settler mortality is a valid instrument. Because this set is evaluated at a single value of $\Sigma$, however, the figure does not account for uncertainty over the reduced form parameters. In contrast, the posterior credible interval for $\rho_{u\zeta}$ in Panel (III) averages both over the posterior for $\Sigma$ and over the conditional identified sets themselves, via a uniform reference prior.\footnote{See (ref).} This interval shows that, averaging over reduced form draws, the relative area of the conditional identified compatible with a valid instrument is very small. Notice the stark contrast between our credible interval for the parameter $\rho_{u\zeta}$ in Panel (III) and that for the identified set for $\rho_{u\zeta}$ in Panel (II). Panel (II) shows that we cannot exclude the possibility that the identified set for $\rho_{u\zeta}$ includes zero, averaged over uncertainty in $\Sigma$. In contrast, Panel (III) shows that one would need to place an inordinate amount of a priori probability over very small regions of the identified set to support the claim that $z$ is a valid instrument.
The primary question of interest, of course, is not the validity of settler mortality as an instrumental variable, but the causal effect of institutions on development. The colored region in Figure (ref) shows how $\kappa$, $\rho_{u\xi^*}$ and $\rho_{u\zeta}$ map into corresponding values for $\beta$. Blue indicates a positive treatment effect, red a negative treatment effect, and white a zero treatment effect. In both directions, darker colors indicate larger magnitudes. As seen from the figure, we cannot rule out negative values for $\beta$. The posterior credible set for the identified set for $\beta$ from columns 3--4 of Panel (II) tells the same story, while accounting for sampling uncertainty in $\Sigma$. Notice from Figure (ref), however, that at least when evaluated at $\widehat{\Sigma}$, the identified set implies negative values for $\beta$ only in the region where $\rho_{u\xi^*}$ is extremely large and there is very little measurement error ($\kappa$ is close to one). Because the posterior for $\underline{\beta}$ is determined entirely from these extreme points, the resulting inference is very conservative, a concern that we raised above in (ref). This observation motivates the idea of averaging not only over reduced form draws $\Sigma$ but also over the conditional identified set itself, as we do in Panel (III), using a uniform reference prior. Unlike the posterior credible interval for the identified set for $\beta$ in Panel (II), our posterior credible interval for the partially identified parameter $\beta$, constructed under a conditionally uniform reference prior, contains only positive values.\footnote{See (ref) for a detailed discussion of the difference between inference for the identified set and inference for the partially identified parameter.} This indicates that the conditional identified sets for $(\kappa, \rho_{u\xi^*}, \rho_{uz})$ contain, on average, only a small region in which $\beta$ is negative.\footnote{Because the prior is uniform, “small” refers to the relative area of a region on the identified set: in Figure (ref), for example, the red region is small compared to the blue and white regions.} Indeed, the posterior median for $\beta$ is 0.49, very close to the OLS estimate from Acemoglu2001. As we see from (ref), the posterior from which the credible interval in Panel (III) was constructed, the IV estimate is very likely an overestimate. In spite of the likely negative correlation between settler mortality and $u$ under reasonable prior beliefs that accord with the data, the main result of Acemoglu2001 continues to hold: it appears that the effect of institutions on income per capita is almost certainly positive.
We now consider an application in which our framework leads to very different conclusions from those of the preceding example. BeckerWoessmann study the long-run effect of the adoption of Protestantism in sixteenth-century Prussia on a number of economic and educational outcomes, using variation across counties in their distance to Wittenberg -- the city where Martin Luther introduced his ideas and preached -- as an instrument for the Protestant share of the population in the 1870s. Here we consider their estimates of the effect of Protestantism on literacy, based on the specification
where $\mathbf{x}$ is a vector of demographic and regional controls.\footnote{In this exercise we include the controls listed in Section III of BeckerWoessmann, specifically: the fraction of the population younger than age 10, of Jews, of females, of individuals born in the municipality, of individuals of Prussian origin, the average household size, log population, population growth in the preceding decade, the fraction of the population with unreported education information, and fraction of the population that was blind, deaf-mute, and insane.}
BeckerWoessmann express beliefs about the three key parameters in our framework. First, their IV strategy relies on the assumption that $\rho_{u\zeta}=0$, an assumption that we will relax below. Second, the authors argue that the 1870 Prussian Census is regarded by historians to be highly accurate. As such, measurement error in the Protestant share should be fairly small. Finally, BeckerWoessmann go through a lengthy discussion of the nature of the endogeneity of the Protestant share, suggesting that it is most likely that Protestantism is negatively correlated with the unobservables:
Results for the “Was Weber wrong?” example appear in Table (ref). Estimates and bounds for $\beta$ indicate the percentage point change in literacy that a county would experience if its share of Protestants were to increase by one percentage point. All other values in the table are unitless: they are either probabilities, correlations, or variance ratios. OLS and IV estimates and standard errors, along with the estimates of the lower bounds $L$ for $\kappa$ appear in row four of Panel (I). Panel (II) presents inference for the identified set. The first column of Panel (II) gives the fraction of posterior draws for the reduced form parameters that yield an empty identified set, while the second column gives the fraction that are compatible with a valid instrument: $\rho_{u\zeta} = 0$. The third and fourth columns of Panel (II) present 90% posterior credible intervals for the identified sets for $\rho_{u\zeta}$ and $\beta$, constructed by symmetrically expanding around the conditional identified set evaluated at the posterior mean for $\Sigma$, as described in (ref). In contrast, panel (III) presents posterior medians and 90% highest posterior density intervals for the partially identified parameters $\rho_{u\zeta}$ and $\beta$.
As we see from Table (ref), BeckerWoessmann obtain an OLS estimate of $0.10$ and an IV estimate that is nearly twice as large: $0.19$ with a standard error of $0.03$. If the instrument is valid, this corresponds to just under a 0.2 percentage point increase in literacy from each percentage point increase in the prevalence of Protestantism in a given county. The estimated lower bound for $\kappa$ in this example is just under a half, which means that at most 50 percent of the measured variation in the Protestant share can be attributed to measurement error. Notice that this bound is somewhat weak: it allows for far more measurement error than one might consider reasonable given the author's arguments concerning the accuracy of the Prussian census data.
Figure (ref) depicts the identified set for $(\kappa, \rho_{u\xi^*}, \rho_{u\zeta})$ evaluated at the posterior mean for $\Sigma$. As above, the surface is colored to indicate the corresponding value of $\beta$: blue indicates a positive treatment effect, red a negative effect, and zero no effect. In both directions, darker colors indicate larger magnitudes. We see immediately from the figure, that unless $\rho_{u\xi^*}$ is large and positive, the treatment effect will be positive, irrespective of the amount of measurement error. The rectangular region surrounded by thick black boundaries indicates our approximation to the prior beliefs of BeckerWoessmann: negative selection, and measurement error that is not too severe. This area is well within the blue region, corresponding to a positive treatment effect. Although it is somewhat harder to see from the figure, the region enclosed in the black boundary also contains $\rho_{u\zeta} = 0$. The belief that $\rho_{u\xi^*}<0$ and measurement error is modest indeed appears to be compatible with a valid instrument in this example.