Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
38,455 characters · 12 sections · 34 citation commands
Possibilistic Instrumental Variable Regression
Instrumental variables (IVs) offer an approach to estimating treatment effects in the presence of unobserved confounding. An IV is an observed variable that must satisfy three key assumptions:
Assumptions A2 and A3 are typically grouped together and referred to as exogeneity or validity of the instruments, while A1 is known as relevance. In practice, it is often difficult to find variables that meet all three criteria, and there can be substantial uncertainty about whether candidate instruments are truly valid.
We introduce a method designed to accommodate slight violations of the exogeneity assumption. If the instruments are not assumed to be valid a priori, there is no one-to-one correspondence between the (always identifiable) reduced-form parameters and the structural parameters of interest. Based on possibility theory zadeh_fuzzy_1999, dubois_possibility_2015, our proposed approach can still perform inference on the structural parameters by assigning the “most-possible” reduced-form parameters. We allow the user to specify a set of potential violations and then compute a conditional posterior distribution of the treatment effect given this set. This will generally be uninformative if the set is “too large”, but can be informative if the violations are small. Unlike many existing methods, our method even supports inference with a single invalid instrument. Based on the possibilistic inferential model construction martin_inferential_2013, martin_possibilistic_2025, we prove a frequentist coverage guarantee for our uncertainty intervals that holds if the violation set is chosen correctly.
An important use case we envision for our method is sensitivity analysis. Classical small_sensitivity_2007, armstrong_sensitivity_2021, cinelli_omitted_2025 as well as Bayesian (or quasi-Bayesian) conley_plausibly_2012, chernozhukov_plausible_2025 proposals exist in the literature. chib_bayesian_2018 and chernozhukov_plausible_2025 use semi-parametric Bayesian methods to analyse a structural model characterized by a set of moment restrictions, while allowing for the possibility that (some of) these moment conditions do not hold exactly. While chib_bayesian_2018 focus only on misspecification of overidentifying restrictions, chernozhukov_plausible_2025 develop a quasi-Bayesian framework to allow for the possibility that all restrictions are invalid. All approaches that we are aware of in the literature are probabilistic and widen uncertainty intervals to reflect the additional uncertainty about the instruments' validity. Our contribution is similar, but has the advantage of doing so naturally, as the widened intervals arise directly from epistemic uncertainty about the parameters rather than from ad-hoc adjustments. The literature on partial identification in instrumental variable models is closely related conley_plausibly_2012, watson_bounding_2024, penn_budgetiv_2025. Our inference is informative only when the invalidity parameter can be constrained to a small set. In that case, all causal effects in the partial identification region have equal posterior possibility. As such, our method emphasises interval estimation rather than point estimation.
The proposed approach is related to the literature on identification and estimation with invalid instruments. Under the plurality rule, the treatment effect can be identified without knowing in advance which instruments are valid kang_instrumental_2016. A sufficient condition is that fewer than half of the instruments are invalid. Estimation in this setting typically relies on $\ell_1$-penalisation kang_instrumental_2016, windmeijer_use_2019 or voting/searching strategies guo_confidence_2018, windmeijer_confidence_2021, guo_causal_2023 to recover the valid instruments. Alternative approaches average across different instrument choices steiner_bayesian_2025 or assume that the direct effects on the outcome and the treatment are orthogonal kolesar_identification_2015. In contrast, our approach is not constrained by such restrictive assumptions and can produce informative—though sometimes very diffuse—results even if only a single invalid instrument is available.
Here, we provide a brief introduction to possibility theory. For more details, we refer to houssineau_parameter_2020, houssineau_robust_2022, hieu_decoupling_2025, martin_possibilistic_2025.
An uncertain variable $\bo{\theta}$ is a mapping from $\Omega_{\mathrm{u}} \to \Theta$, where $\Omega_{\mathrm{u}}$ is a sample space of deterministic phenomena. We think of $\Omega_{\mathrm{u}}$ as containing a true (but unknown) reference element $\omega_{\mathrm{u}}^*$, thus there is no aleatoric uncertainty connected to $\Omega_{\mathrm{u}}$. We characterise the uncertain variable $\bo{\theta}$ by a possibility function $f_{\bo{\theta}} : \Theta \to [0, 1]$ with $\sup_{\theta \in \Theta} f_{\bo{\theta}}(\theta) = 1$. This possibility function gives rise to an outer probability measure
for a subset $A \subseteq \Theta$. Unlike a regular probability measure, this outer measure is not additive with respect to disjoint sets. In fact, the outer measure can be seen as an upper bound on probability measures. Thus, the possibility $\Bar{\mathbb{P}}_{\bo{\theta}}(A)$ is the maximum subjective probability one would be willing to assign to the set $A$. The possibility function $f_{\bo{\theta}}(\theta) = 1$ is the most uninformative possibility function in the sense that it assigns full credibility to any (non-empty) set.
Let $\bo{\psi}$ be another uncertain variable on $\Psi$ such that $\bo{\theta}$ and $\bo\psi$ have joint outer measure $$\bar{\mathbb{P}}_{\bo\theta, \bo\psi}(A \times B) = \sup_{\theta \in A, \psi \in B} f_{\bo\theta, \bo\psi}(\theta, \psi), \quad B \subseteq \Psi,$$ where $f_{\bo\theta, \bo\psi}$ is a joint possibility function. Marginalising over $\theta$ is done by setting $A = \Theta$ such that the marginal possibility function of $\bo\psi$ is $f_{\bo\psi}(\psi) = \sup_{\theta \in \Theta} f_{\bo\theta, \bo\psi}(\theta, \psi)$. Conditional outer measures can be defined analogously to probability theory as
such that $f_{\bo\theta \mid \bo\psi}(\theta \mid \psi) = f_{\bo\theta,\bo\psi}(\theta, \psi) / f_{\bo\psi}(\psi)$ for all $\psi \in \Psi$ with $f_{\bo\psi}(\psi) > 0$. If $\bo\psi = T(\bo\theta)$ is a transformation of $\bo\theta$, we have that
where the appropriate convention is $\sup \emptyset = 0$. There is no need to account for the change in measure by a Jacobian term. This difference to probability theory plays an important role in our proposed methodology.
In this paper, we propose to do Bayesian inference with possibilistic priors. Consider the random variable $Y$ characterised by the statistical model $\{P_\theta : \theta \in \Theta \}$ with corresponding probability density $p(\cdot \mid \theta)$. We incorporate prior information on the parameter $\theta$ in the form of a possibility function $f_{\bo\theta}$. Then, the posterior possibility function is
The main differences from standard Bayesian inference are that the prior is represented by a possibility function and the denominator is based on maximisation rather than integration. These differences are small enough that much of the intuition from standard Bayesian inference carries through. The possibilistic framework allows vacuous prior information to be modeled by an uninformative possibility function, whereas in standard Bayesian inference an improper prior may yield an improper posterior. It also provides clear computational benefits, since optimisation is generally no harder than integration.
Let $Y_i$ denote the outcome of interest, $X_i$ a treatment or endogenous variable, and $Z_i$ a $p$-dimensional (row) vector of instrumental variables. We observe $n$ i.i.d. copies of $\{Y_i, X_i, Z_i\}_{i=1}^n$ generated from the structural model
where the errors are assumed to be jointly Gaussian, $(\epsilon_i, \eta_i)^\intercal \sim N(0, \Sigma)$. Whenever $\Sigma$ is not diagonal, this indicates unobserved confounding (or endogeneity), and “naive” inference in the outcome model delivers biased results. In this setting, the instruments $Z_i$ are relevant if $\gamma_2 \neq 0_p$ and exogenous if $\alpha = 0_p$. Rather than enforcing these assumptions a priori, we incorporate uncertainty about them into the model. In particular, our approach allows for $\alpha$ to be non-zero.
Consider the “reduced-form” equation model
where $\gamma_1 = \beta \gamma_2 + \alpha$ and $\Psi$ is the reduced-form covariance given by $$\Psi = R(\beta) \Sigma R(\beta)^\intercal , \quad R(\beta) =
.$$ Equivalently, we have that the matrix $W =
$ of stacked outcomes and treatments follows a matrix Normal distribution, $W \sim MN(Z \Gamma, I_n, \Psi)$, where $Z$ is the matrix with $i$-th row $Z_i$, and the coefficient matrix is $\Gamma =
$.
The structural parameters are identifiable if we can find a unique solution $(\alpha, \beta, \gamma_2, \Sigma)$ given $(\gamma_1, \gamma_2, \Psi)$, or equivalently, if there exists a bijective mapping between the reduced-form and the structural parameters. The matrix $R(\beta)$ is invertible for any $\beta \in \mathbb{R}$ and $\gamma_2$ maps to itself. Thus, the structural parameters are identifiable if and only if
has a unique solution $(\alpha, \beta)$ given $\Gamma$. Without any assumptions on $\alpha$, this is not the case. Typically, one assumes $\alpha = 0_p$, or that at least the majority of its components are zero kang_instrumental_2016.
We propose to perform possibilistic posterior inference in the reduced-form model and propagate the uncertainty to the structural parameters. Let $f$ be a prior possibility function on $(\Gamma, \Psi)$, then the posterior possibility $f_{\mathrm{RF}}$ is given by
where $\mathbb{S}^2_{+}$ is the cone of positive-semidefinite and symmetric $2 \times 2$ matrices. Under the uninformative prior $f(\Gamma, \Psi)=1$, the supremum in the denominator is attained by the standard maximum-likelihood estimators.
Then we can define a possibility function $f_{\mathrm{S}}$ for the structural parameters as
This operation is not well-defined for a probabilistic posterior distribution, as mapping the reduced-form to the structural parameters is not one-to-one in general. Under uninformative prior possibility functions on the reduced-form parameters, this optimisation problem can be solved in closed form (see Appendix (ref)). We have that \( f_{\mathrm{S}}(\alpha, \beta, \Sigma \mid W) = f_{\mathrm{RF}}(\Gamma^*(\alpha, \beta, \Sigma), R(\beta) \Sigma R(\beta)^\intercal \mid W) \), where the optimal reduced-form coefficient matrix is \[ \Gamma^*(\alpha, \beta, \Sigma) = \Hat{\Gamma} + \frac{1}{\sigma_{11}} \left(\alpha - \Hat{\Gamma}
\right)
\Sigma R(\beta)^\intercal, \] with \( \Hat{\Gamma} = (Z^\intercal Z)^{-1} Z^\intercal W \) denoting the least-squares estimate of the reduced-form coefficient matrix, and \( \sigma_{11} \) representing the marginal variance of the outcome in the structural model.
Our main object of interest is the posterior possibility of $\beta$ given that $\alpha$ lies in a violation set $A$. The following proposition characterises this conditional posterior possibility function.
The extreme case of $A = \mathbb{R}^p$, where $\alpha$ is completely unconstrained, results in the uninformative marginal possibility function of $\beta$, that is for all $\beta \in \mathbb{R}$ we have that $$f(\beta \mid W) = f(\beta \mid \alpha \in \mathbb{R}^p, W) = 1.$$ This is not surprising, as $\beta$ is not identified without extra information on $\alpha$, thus we get back the prior.
More generally, the intersection of the affine subspace \(t(\beta)\) and the violation set \(A\) defines the partial identification region, in which all values of \(\beta\) are equally plausible. To obtain informative results, \(A\) must be sufficiently restrictive so that, for most values of \(\beta\), the implied \(\alpha\) lies outside this region.
If $A$ is a rectangle, computing the projection is a standard quadratic programming problem. Alternatively, we can also bound a norm of $\alpha$, that is specify the constraint set as $A_\tau = \left\{ \alpha : \lVert \alpha \rVert \leq \tau \right\}$, where the threshold $\tau$ is the maximum invalidity budget across all instruments penn_budgetiv_2025. More details are provided in Appendix (ref). In either case, it is important that the choice of $A$ corresponds to the scale of the instrument data. To simplify the interpretation, it may be useful to standardise the instruments.
Our procedure can be viewed from two different perspectives. The first treats the choice of \(A\) as partial prior information based on domain knowledge, specified by the analyst before observing the data. Incorporating this prior information through post-hoc conditioning, rather than directly specifying a prior possibility function, is convenient for two reasons: (1) it allows closed-form solutions for many of the expressions, and (2) specifying a prior on the reduced-form parameters that induces the desired prior on the structural parameters is challenging. The second perspective emphasises sensitivity analysis for a particular effect, where the analyst gradually widens \(A\) to assess how much instrument invalidity would be required for the effect to disappear.
The sampling properties of our posterior possibility can be improved by using the validification procedure proposed by martin_inferential_2013. Specifically, we transform the posterior possibility defined in Proposition (ref) to the validified posterior possibility function
where \(w\) denotes the observed value of \(W\), and \(P_\beta\) represents the probability measure of \(W\) as a function of \(\beta\). The following proposition shows that the validified posterior possibility controls the type-I error, that is, its probability of assigning “too little” possibility to the true value of $\beta$ is bounded at the nominal level, as long as the violation set $A$ contains the true value of $\alpha$.
An immediate corollary is that the upper level sets are valid confidence sets.
These results guarantee that our inference controls the type-I error if $A$ contains the true value of $\alpha$. However, at the same time we want to maximise efficiency (minimise the type-II error), which is inversely proportional to the volume of $A$. Thus, one faces a trade-off between choosing $A$ large enough for it to likely contain the true $\alpha$, but not choosing it too large so that the inference becomes overly conservative. In practice, this choice should be based on domain knowledge. When sensitivity analysis is the goal, one may be interested in finding the largest set $A$ such that a particular effect holds (e.g., such that $0 \notin C_\delta(w, A)$).
The validified posterior possibility \(\pi_w(\cdot \mid A)\) is not available in closed form, but a natural Monte Carlo approximation is
where $W_i$ are independent samples from $P_\beta$. This approximation offers exact results (up to Monte Carlo error), but is computationally expensive. A cheaper approximation is the Wilk's style $\chi^2$ approximation $$\pi_w(\beta \mid A) \approx 1 - F(-2 \log f(\beta \mid \alpha \in A, W = w)),$$where $F$ is the cdf of a $\chi^2$ random variable with $1$ degree of freedom. This approximation is based on asymptotic Gaussianity and is less appropriate whenever the parameter is not point-identified, i.e., when $A$ is not a singleton. For a comprehensive treatment, we refer to martin_possibilistic_2025. The case with non-vacuous prior information is covered in martin_valid_2023, where the probability measure $P_\beta$ has to be replaced by a suitable outer measure.
First, we consider a toy example with a single instrument and varying instrument validity. The single instrument is generated from a standard Gaussian distribution. The residual pairs are simulated from a bivariate Gaussian with unit variances and correlation $\rho = 1/2$. Then, we generate pairs $(Y_i, X_i)$ from ((ref)) with $\gamma_2 = 1, \beta = 1$, and varying invalidity parameter $\alpha \in \{0, 0.25, 0.5\}$. The outcome and treatment are centred so that we do not need to include an intercept.
Our performance criterion is the empirical coverage of a $95\%$ uncertainty interval for the treatment effect $\beta$. We consider both the $\chi^2$ approximation and Monte Carlo (MC) sampling from the validified posterior possibility under different tolerated sets $A$. We compare our empirical coverage to those of naive two-stage least squares (TSLS), plausible generalised method of moments with a Gaussian prior (PGMM-g) chernozhukov_plausible_2025, and BudgetIV penn_budgetiv_2025. The results are given in Table (ref) and the full details are provided in Appendix (ref). Code to reproduce our findings is available at https://github.com/gregorsteiner/PossibilisticIV.
For $A = \{0\}$, coverage falls as the true $\alpha$ moves away from zero. Widening $A$ to include the true $\alpha$ maintains coverage above the nominal level, but overly wide $A$ makes the intervals excessively conservative. The Monte Carlo-based possibility functions tend to be slightly better than the $\chi^2$ approximation in this setting. We believe the slight undercoverage for $A = \{0\}$ when $\alpha =0$ is due to Monte Carlo error. The only other method that maintains coverage above the nominal level is BudgetIV with a budget of $0.5$. However, it is overly conservative even when the true $\alpha = 0.5$.
In the second experiment, we consider $p=5$ instruments generated from a multivariate standard Gaussian, $s$ of which are invalid. The coefficient $\alpha$ is chosen such that the first $s$ components are $0.1$ and the remaining $p-s$ components are zero. We set $\gamma_2 = (1/4, \ldots,1/4)$, which corresponds to a first-stage $R^2$ of approximately $1/4$, thus the instrument strength is moderate. The treatment effect is again set to $\beta = 1$ and the residuals are generated as above. Now, we also include the confidence interval method (CIIV) windmeijer_confidence_2021 and gIVBMA steiner_bayesian_2025.
Table (ref) shows the results. For $s = 0$, all methods achieve good coverage. As more instruments become invalid, naive approaches, including possibilistic IV with $A = \{0\}$, lose coverage. Our two variants with $A$ containing the true $\alpha$ maintain good coverage even when all instruments are invalid. However, when $\alpha$ lies on a corner of the hypercube (e.g., $A = [0.0, 0.2]^p$ with $s = 0$ or $A = [-0.1, 0.1]^p$ with $s = 5$), the intervals can be slightly too short, particularly for the $\chi^2$ approximation. BudgetIV is the only other method that can maintain coverage above the nominal level for $s=5$.
Our approach widens the confidence sets appropriately when the instruments are not assumed to be valid a priori. This ensures valid inference even if no valid instruments exist. These sets may include a non-unique mode, reflecting partial identification. However, as shown in the experiments above, choosing $A$ too wide can make the confidence sets overly conservative and less useful in practice. BudgetIV can similarly maintain good coverage, as it also relies on partial identification, yet it may not return a plausible set for certain hyperparameter specifications (see Appendix (ref)). In contrast, our method always yields a posterior possibility function, even if it is sometimes very uninformative.
We illustrate our method with an empirical application estimating the effect of institutions on economic output. The analysis uses the dataset of 64 countries originally compiled by acemoglu_colonial_2001 and reexamined by chernozhukov_plausible_2025 to demonstrate their quasi-Bayesian approach that allows for small violations of instrument exogeneity.
The outcome is log GDP per capita in 1995, with the main predictor being protection against expropriation, a proxy for institutional quality. A central challenge is endogeneity: institutions may affect income, but income may also influence institutions. To address this, acemoglu_colonial_2001 use the mortality of early European settlers as an instrument, arguing that long-term settlement incentives shaped institutional quality. The instrument’s validity rests on the assumption that settler mortality is exogenous (conditional on covariates). Following chernozhukov_plausible_2025, a direct effect of settler mortality should not be stronger than $0.1$ in absolute terms, if any, justifying $A = [-0.1, 0.1]$ as a reasonable violation set. We consider the specification with an intercept and (normalised) distance from the equator as exogenous control variables and (log) settler mortality as the sole instrument. We project out the covariates and run our analysis on the residuals resulting from that projection.
Figure (ref) shows validified posterior possibility functions for the valid case and allowing some invalidity. The curve corresponding to $A = [-0.1, 0.1]$ is wider and lacks a unique mode, reflecting the partial identification of \(\beta\). Both indicate a significantly positive effect with heavier right tails, consistent with chernozhukov_plausible_2025. For comparison, we also include the violation set $A = [-0.4, 0.4]$, which is wide enough for the significant effect to disappear. Table (ref) in Appendix (ref) shows that our uncertainty intervals are narrower than those of chernozhukov_plausible_2025, especially on the right tail. Relaxing the exogeneity assumption and allowing for $\alpha \in [-0.1, 0.1]$ does not qualitatively change the conclusion that good institutions promote economic output.
For $A = \{0\}$, the $\chi^2$ and MC approximations look essentially identical. For larger $A$, however, the latter display steeper decay away from the partial identification region. This is expected as the Gaussian approximation cannot correctly capture the tail behaviour when $\beta$ is not point-identified. In these cases, we observe that the $\chi^2$ approximation leads to more conservative (but still valid) inference. This suggests a practical strategy can be fitting the $\chi^2$ approximation first, and then only fitting the computationally expensive MC approximation if additional tightness is required.
To further interpret the obtained posterior possibility, we can view it as an upper bound on a precise (subjective) probability measure. This allows to deduce a corresponding lower bound, and therefore provides a probability interval for the event of interest. Table (ref) in the Appendix displays such intervals for the hypothesis $\beta > 0$, given by the pair $\left( 1 - \sup_{\beta \leq 0} \pi_w(\beta \mid A), \sup_{\beta > 0} \pi_w(\beta \mid A) \right)$, under different choices of the violation set $A$. It takes $\alpha$ being close to $0.4$ in absolute terms to materially change the inference, that is, the lower probability drops below the nominal level. Thus, the qualitative effect of institutions on economic output is very robust to reasonable violations of the exogeneity assumptions.
We revisit the relationship between education and wages using the dataset from card_using_1995. Education choices are not random, but are influenced by unobserved traits that also affect earnings, so the observed link between schooling and wages may not reflect the true causal effect. card_using_1995 adresses this endogeneity by using college proximity as an instrumental variable. The outcome is (the logarithm of) hourly wages and the treatment is years of schooling received, which we instrument with proximity to a four-year college. We partial out the available covariates, including experience, experience squared, family background, marital status, race, region, and parents’ education. The parental education variables are missing for a substantial proportion of the sample. Following card_using_1995, we impute them using the mean and include an indicator for missingness.
We also expect the direct effect of proximity to a college on hourly wages to be positive, if it exists. A magnitude of the effect in between zero and $4\%$ seems reasonable to us. Thus, we consider the violation sets $A = [0, 0.02]$ and $A = [0, 0.04]$. Figure (ref) presents the results. For $A = \{0\}$, the posterior mode is just below $0.15$, comparable to the estimates in card_using_1995, and the effect is significant, since the uncertainty intervals do not include zero. In the other two cases, the left tail is shifted to the left and the effect is no longer significant, since the uncertainty intervals now include zero. Thus, positive returns to education are rather sensitive to violations of the exogeneity assumption of the instrument “college proximity”. Finally, under the MC approximation, the right tail is thinner, assigning lower possibility to very high returns to schooling.
As in the previous example, we can compute lower and upper probabilities for the hypothesis that the returns to schooling are positive, see Table (ref) in the Appendix. We see that the positive treatment effect is robust to a direct effect of college proximity on wages of $1\%$, whereas larger violations lead the lower probability to drop significantly. For a direct effect of $4\%$, the lower probability of positive returns to schooling is only approximately $0.52$.
In this paper, we propose a method for instrumental variable regression based on possibility theory. This allows for valid inference under some user-specified violations of the instrument exogeneity assumption. The proposed method also provides a very natural way to perform sensitivity analyses. The resulting inference is based on first principles and directly reflects the epistemic uncertainty about the parameters. Our approach performs well in simulation experiments and delivers credible results in real-data applications.
Unlike some alternative approaches, our method does not infer the instrument validity. Instead, inference is performed only relative to a user-defined set of possible violations. As a result, the effectiveness of the method depends on the user’s knowledge about which instruments might be invalid and the nature of their potential violations. If the specified set of violations is too large, the resulting inference will tend to be overly conservative, or even completely uninformative.
We thank Ryan Martin for helpful comments that substantially improved this paper. This research is supported by the Ministry of Education, Singapore, under its Academic Research Fund Tier 1 (RS02/24), and by the Singapore Ministry of Digital Development and Information under the AI Visiting Professorship Programme (AIVP-2024-004).