EconBase
← Back to paper

Robust Bayesian Method for Refutable Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,606 characters · 25 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Bayesian Method for Refutable Models

\affil[1]{The University of Sydney}

abstract\begin{spacing}{1.2} { We propose a robust Bayesian method for economic models that can be rejected by some data distributions. The econometrician starts with a refutable structural assumption which can be written as the intersection of several assumptions. To avoid the assumption refutable, the econometrician first takes a stance on which assumption $j$ will be relaxed and considers a function $m_j$ that measures the deviation from the assumption $j$. She then specifies a set of prior beliefs $\Pi_s$ whose elements share the same marginal distribution $\pi_{m_j}$ which measures the likelihood of deviations from assumption $j$. Compared to the standard Bayesian method that specifies a single prior, the robust Bayesian method allows the econometrician to take a stance only on the likeliness of violation of assumption $j$ while leaving other features of the model unspecified. We show that many frequentist approaches to relax refutable assumptions are equivalent to particular choices of robust Bayesian prior sets, and thus we give a Bayesian interpretation to the frequentist methods. We use the local average treatment effect ($LATE$) in the potential outcome framework as the leading illustrating example. } \end{spacing} {Keywords---} Refutable Models; Non-parametric Bayesian; Robust Bayesian; LATE {JEL---} C12, C13, C18, C51, C52

Introduction

Model assumptions are often crucial for identification in many econometric models. Often, applied researchers use convenient econometric assumptions to derive an informative identified set of parameters of interest. However, assumptions may be rejected by some data distribution, which leads to an empty identified set. When such a problem arise, we call the assumption refutable. To facilitate the discussion, suppose that the original refutable assumption $A$ can be written as the intersection of several weaker assumptions, i.e., $A=\cap_{l} A_l$, and the econometrician knows that $A_j$ will lead to the refutation issue.

When an economic assumption is refuted, there are several ways to salvage the refutable assumptions. Frequentist has two major ways to deal with a refutable assumption: The first approach is to completely give up $A_j$ that can lead to data rejection and then identify the parameter of interest under the weaker assumption. However, this approach often gives up the original economic theory behind the assumption. As a result, we may get a wide identified set that is not informative of the parameter of interest. Another approach is to minimally relax $A_j$ until the relaxed assumption cannot be rejected and then identify the parameter of interest under the minimally relaxed assumption. This approach is similar to the sensitivity analysis. However, it is hard to justify the rationality of the minimally relaxed assumption and the interpretation of the consequent indentified set becomes vague.

In contrast, the Bayesian approach of dealing with refutable assumption can deliver better interpretation of the relaxed assumption. The Bayesian econometrician first specifies a prior belief over all econometric structures. To incorporate the refutable assumption as a reasonable economic assumption, the econometrician may specify a prior belief with a higher prior probability on the original assumption. However, specifying a unique prior belief requires the econometrician to take a prior stance not only on the likelihood of violation of $A_j$, but also on all other aspects of the model. As a result, specifying a unique prior belief may overstate the econometrician's belief over all possible econometric structures.

This paper proposes using a robust Bayesian method to deal with refutable assumptions. Compared to the standard Bayesian method, our robust Bayesian method allows the econometrician to consider a set of priors that has a common marginal distribution. This common marginal distribution measures the econometrician's belief of how likely $A_j$ is violated. As a result, the econometrician can choose to only take a stance on devation from $A_j$ while remain agnostic about other aspects of the econometric model. There are several appealing features of our robust Bayesian method. First, it can be shown that under proper choices of the common marginal distribution, the robust Bayesian method is equivalent to certain frequentist approaches to salvage the refutable model. More specifically, the robust Bayesian method estimated identified set will converge to the frequentist identified set.\footnote{We do not seek the equivalence result for inference because many econometric models are nonparametric.} As a result, we can interpret the frequentist approaches to the refutable models as a choice of the robust Bayesian prior set. Second, unlike the standard Bayesian method which often generates a point estimate from the posterior, the robust Bayesian method can generate a posterior mean set which is similar to the partial identification method. Third, in many nonparametric models, the common marginal constraint allows us to find computationally feasible ways to compute the posterior mean set.

To preview the robust Bayesian framework, we first formalize some essential notations in the previous discussion: Recall that the original refutable assumption $A$ as the countable intersection of assumptions $A=\cap_{l=1}^\infty A_l$, where each assumption $A_l$ is a collection of econometric structures $s$. Each econometric structure $s$ specifies an economic mechanism and predicts a unique data distribution $F$. The Bayesian econometrician takes a stance on which assumption $A_j$ is likely to be violated and only relaxes $A_j$. For the assumption $A_j$ that will be relaxed, the econometrician choose a distance function $m_j(s')$ that measures the distance of any econometric structure $s'$ to $A_j$. Since each econometric structure $s$ predicts a unique data distribution, each prior belief $\pi_s$ over structures induce a joint distribution of $(s,F,m_j(s))$. To avoid over-specification of her belief, the econometrician considers a particular set of beliefs $\Pi_s$ whose elements are supported on $\cap_{l\ne j} A_l$ with the constraint that all $\pi_s\in \Pi_s$ induce the same conditional marginal conditional marginal distribution $\pi^*_{m_j|F}$. By choosing to further condition $\pi^*_{m_j|F}$ on $F$, we can allow more flexible characterization of the prior beliefs and we can show some equivalence results to the frequentist-based approaches.

After observing realized data $\bm{X}_n$, the econometrician updates each belief in the prior set and collect the posteriors as $\Pi_{s|\bm{X}_n}$. We follow giacomini2021robust to study two statistics from the posterior set. The first one is posterior-mean bound. For each posterior belief in the posterior set, we can derive the posterior mean of the parameter of interest. By maximizing and minimizing the posterior mean of the parameter of interest, we can derive a bound. It can be shown that by choosing $\pi^*_{m_j|F}$ to be a proper degenerate marginal distribution, the posterior mean bound is the convex hull of the identified set under the frequentists' minimally relaxed assumption approach. Characterizing the maximum and minimum can be hard because we are maximizing over a set of posterior distributions. We then show that we can transform the optimization problem to an optimization over all structures with a fixed amount of deviation, i.e. $m_j(s)=m$, and then integrate a corresponding parameter quantity with respect to the joint posterior belief over $(m_j(s),F)$. The second quantity is a confidence interval that covers the parameter of interest with $1-\alpha$ probability under all posterior beliefs. The calculation of the confidence interval is similar to that of the posterior-mean bound, which uses the quantiles of the parameter of interest under the joint posterior belief on $(m_j(s),F)$.

For the main application, we look at the local average treatment effect ($LATE$) in the potential outcome framework. kitagawa2015 and mourifie2017testing show the `No Defiers' assumption and the independent IV assumption are jointly refutable. We proceed to consider a prior belief set with marginal density constraints on the measure of defiers. Since each econometric structure in the $LATE$ framework contains a distribution of potential outcomes which is infinite dimensional, the optimization to find the posterior mean bound of $LATE$ is an infinite dimensional optimization problem. However, we use several tricks to convert the infinitely dimensional optimization problem to a combination of a feasible finite dimensional optimization problem and a integratoin problem. We apply the robust Bayesian method to study the return of college education card1993using and compare our result with that of frequentists' approach. In addition to the main application, we also look at the applications to monotone IV model, the intersection bounds model and the Logit discrete choice model.

The framework in this paper is mostly related to the robust Bayesian method in giacomini2021robust but we add discussion to their paper in several ways. First, giacomini2021robust show that the posterior mean set is equivalent to the frequentist-based partially identified set and the variation of the prior set in their framework does not matter asymptotically. In particular, when the model is parametric, they establish a Berstein-von Mises result for the robust Bayesian method.\footnote{In principle, if only convergence is required, their method can accomodate nonparametric models.} In contrast, we use the robust Bayesian method as a modeling choice to deal with refutable assumptions. More precisely, the choice of the marginal distribution of $m_j$ matters for the final posterior mean bound and the confidence set.\footnote{From this perspective, our method has a Bayesian interpretation because the econometrician has to take a stance on which assumption can fail and how severely the assumption might be violated.} We show some equivalence to some frequentist-based approaches, but we only aim to show that these frequentist-based approaches have a Bayesian interpretation. Second, it is computationally hard to accomodate nonparametric models in giacomini2021robust because of the specific choice of the prior set. In contrast, by further imposing the unique marginal constraint $\pi_{m_j|F}^*$, we can simplify the computation in many econometric settings. However, we do not establish the nonparametric Berstein-von Mises result.

\paragraph{Related Literature} Our paper contributes to the literature on salvaging refutable or misspecified models. masten2018 propose an ex-post way to salvage refutable models by considering a minimal relaxation method. bonhomme2018minimizing consider a local sensitivity analysis of a potentially misspecified model. christensen2023counterfactual consider a sensitivity analysis of the counterfactuals to parametric model assumptions.

We also contribute to the use of the robust Bayesian method. giacomini2021robust show the partial-identification version of Bernstein-von-Mises equivalence theorem for parametric model. The robust Bayesian method is also applied to study the uncertain identification in SVARs model giacomini2022robust. In statistics, the robust Bayesian method has been studied by deroberts1981bayesian, berger1986robust, wasserman1989robust, wasserman1992computing.

Our application to the $LATE$ framework contributes to the study of $LATE$ under the `No Defiers' condition. { Since kitagawa2015 proves the sharp testable implication of IA1994, literature relaxes the `No Defiers' condition. chaisemartin2018Defier discusses the economic meaning of the conventional $LATE^{Wald}$ when there are defiers. He shows that the $LATE^{Wald}$ identifies the net average treatment effect of a subgroup of compliers after deducting the average treatment effect of defiers.}

The rest of the paper is organized as follows. Section (ref) uses the potential outcome model IA1994 to introduce the robust Bayesian method. We use this section as a concrete example to illustrate the robust Bayesian method procedures and also the possible difficulties. Section (ref) describes a theory of robust Bayesian method. We also show the general frequentist-equivalence result in this section. Section (ref) considers further applications. Section (ref) concludes. All proofs are collected in Appendices.

A Leading Example of the Potential Outcome Model

We start with an example of a binary treatment and a binary instrument IA1994. This application illustrates the construction of the robust Bayesian prior set and the connnection to frequentists' approach. Notations are set up to match the general theory in the later sections.

The Potential Outcome Framework

An econometrician observes an outcome variable $Y_i$, a treatment decision $D_i$, and a binary instrument $Z_i$. The observed variables $(Y_i,D_i)$ are generated through the following potential outcome framework:

align[align omitted — 133 chars of source]

where $D_i(1),D_i(0)$ are potential treatment decisions, $Y_i(1),Y_i(0)$ are the potential outcomes and $Z_i$ is a binary instrument. We implicitly impose the exclusion restriction in ((ref)).

Variables in ((ref)) can be classified into two types. The variables $\epsilon_i=(D_i(1), D_i(0), Y_i(0), Y_i(1), Z_i)$ reflect the fundamental heterogeneity of the economic agent $i$ and we call them the underlying variables. The variables $X_i=(Y_i, D_i, Z_i)$ are observed variables. Let $\mathcal{Y}$ be the space of $Y_i$ and let $\mathcal{B}$ be the Borel-sigma algebra on $\mathcal{Y}$. We consider two distribution spaces: The space of distributions of $X_i$ is

equation[equation omitted — 123 chars of source]

and space of distributions of underlying variables is

equation[equation omitted — 171 chars of source]

Starting from a $G^s(\epsilon)$, the potential outcome equation ((ref)) defines a unique distribution of observable via the map $M:\mathcal{G}\rightarrow \mathcal{F}$ such that:

equation[equation omitted — 223 chars of source]

In other words, $F$ is the push-forward distribution of $G^s$ under the mapping ((ref)). We call a pair $(M,G^s)$ an econometric structure. In contrast to the $G^s$ that describes the individual heterogeneity, the mapping $M$ describes the economic machinsm to generate observed varaibles. While $M$ is fixed for the potential outcome model, it can vary in the later general theory section.\footnote{Think of an OLS regression $Y_i=\beta_0+\beta_1 X_i+\epsilon_i$, the mapping $M$ depends on the parameter $(\beta_0,\beta_1)$.} We consider a paradigm for analysis of the model:

equation[equation omitted — 176 chars of source]

Empirical researchers often use the Imbens-Angrist Monotonicity assumption (IA-M) where the exogeneity and monotonicity of the instrument $Z_i$ are assumed. We formalize the IA-M assumption (denoted by $A$ ):

equation[equation omitted — 231 chars of source]

where $A^{ND}$ is the `No Defiers' assumption and $A^{IV}$ is the independent IV assumption. The main parameter of interest is the local average treatment effect for compliers:

equation[equation omitted — 95 chars of source]

The IA-M assumption is preferred by applied econometricians as the economic intuition behind it is clear: Defiers are abnormal and assumed away from the model. If the IA-M assumption holds, we can identify the $LATE$ as the Wald ratio IA1994: \[ LATE^{Wald}(F)=\frac{E_F[Y_i|Z_i=1]-E_F[Y_i|Z_i=0]}{E_F[D_i|Z_i=1]-E_F[D_i|Z_i=0]}. \]

The sharp testable implication of $A$

The IA-M assumption can be rejected by some data distributions. We summarize the sharp testable implications in kitagawa2015 who define the following two quantities for all $B\in \mathcal{B}$ and $d\in \{0,1\}$:

equation[equation omitted — 191 chars of source]
lemLet $P_F(\cdot,d)$ and $Q_F(\cdot,d)$, $d\in\{0,1\}$, be absolutely continuous with respect to some measure $\mu_F$. For any structure $s\in A$, $F=M(G^s)$, and any Borel set $B$, the $F$ must satisfy: \begin{equation} \begin{split} P_F(B,1)&\ge Q_F(B,1),\\ Q_F(B,0)&\ge P_F(B,0). \end{split} \end{equation} Moreover, for any $F$ satisfying ((ref)), there is an $s\in A$ such that $F=M(G^s)$.

An implication of ((ref)) is that $E_{F}[D_i|Z_i=1]\ge E_{F}[D_i|Z_i=0]$ must hold. Whenever (ref) fails, the identified $LATE$ is an empty set.

Frequentist approach to the refutable IA-M assumption

There are two approaches used by frequentists to address the issues when the baseline assumption is refutable. We focus on relaxing the `No Defiers' assumption as the illustration. The first approach is to interpret the identified expression under the relaxed assumption. In the $LATE$ example, researchers give up the `No Defiers' assumption and interpret the Wald ratio $LATE^{Wald}$ as compliers' treatment effect net of defiers' treatment effect chaisemartin2018Defier. The second approach is to take an adaptive method to relax the assumption but still identify the average treatment for the compliers dahl2023nevertoolate,liao2020estimating. In this approach, the econometrician first uses the data to select a model\footnote{In particular, the model with a minimal amount of defiers that can rationalize the data distribution. } that permits the prescence of defiers, and then estimate the treatment effect for compliers.

If we take the first approach and give up the `No Defiers' assumption, we also lose the interpretation of the local average treatment for compliers. By giving up the `No Defiers' assumption completely, the identified set of the $LATE$ for compilers is generally unbounded, even if our data distribution passes the test implied by Lemma (ref). By doing so, the econometrician gains the maximal robustness for the presence of defiers, but she also ignores the economic intuition behind the `No Defiers' assumption.

The adaptive approach to relax the assumption is appealing since it keeps the economic rationales of the `No Defiers' assumption and let the data tells the probability of defiers. In particular, when the data distribution passes the test implied by Lemma (ref), the data-selected model will identify the same $LATE$ quantity as the Wald ratio. However, the selection criteria are arbitrary and it may be hard to justify the choice of minimal defier with a consistent statistical framework.\footnote{liao2020estimating interprets the selection criteria as a mixture of Bayesian model selection and a subsequent frequentist estimation stage. }

The Robust Bayesian Approach

An alternative representation of the IA-M assumption

Since the `No Defiers' assumption can lead to testable implications, we may want to relax $A^{ND}$ while keep the $A^{IV}$ assumption. In view of the Bayesian method, we may want to put a prior belief $\pi_s$ supported on $A^{IV}$. However, such a prior $\pi_s$ can have a critical issue, i.e., the $\pi_s$-induced $\pi_F$ is not supported on the whole $\mathcal{F}$: Let $\pi_F$ be the $\pi_s$-induced belief on $\mathcal{F}$ such that for any $\mathcal{F}_1\subseteq\mathcal{F}$, $\pi_F(\mathcal{F}_1)=\int_{\mathcal{S}} \mathbbm{1}(M(G^s)\in \mathcal{F}_1)d\pi_s$, then there exists a non-trivial\footnote{We illustrate the non-trivial subset via a simple example in liao2020estimating. Suppose $Y_i\in\{0,1\}$ is also binary, then $P_F(Y_i=1,D_i=0|Z_i=0)\ge P_F(Y_i=1,D_i=0|Z_i=1)-P_F(D_i=1|Z_i=0)$ must hold for the observed data distribution. Now, since $Y_i$ is binary, we can characterize $F$ using an 8-d vector. It is easy to see that the above testable implication fails in a Lebesgue-positvely-measured set. } subset of $\mathcal{F}'\subseteq \mathcal{F}$ such that $\pi_F(\mathcal{F}')=0$. This problem arises because the $A^{IV}$ still generates testable implications kitagawa2009identification.

Instead, we consider an alternative characterization of the IA-M assumption in liao2020estimating, under which, when we give up the $A^{ND}$, is non-refutable.

lemThe IA-M assumption defined in ((ref)) can be equivalently written as the intersection: $A= A^{TI}\cap A^{EM-NTAT}\cap A^{ND}$ where: \begin{enumerate} • $A^{TI}=\left\{s\big|\, Z_i\perp \left(Y_i(0),Y_i(1)\right)|D_i(1),D_i(0) \right\}$ is the conditional type independent instrument assumption; • Assumption $A^{EM-NTAT}$ is the set of structures $s$ such that: \[ \begin{split} E_{G^s}[\mathbbm{1}(D_i(1)=D_i(0)=1)|Z_i=1]&=E_{G^s}[\mathbbm{1}(D_i(1)=D_i(0)=1)|Z_i=0],\\ E_{G^s}[\mathbbm{1}(D_i(1)=D_i(0)=0)|Z_i=1]&=E_{G^s}[\mathbbm{1}(D_i(1)=D_i(0)=0)|Z_i=0]. \end{split}\] This assumption says that the measure of always/never takers is independent of the instrument. \end{enumerate}

We delegate further discussions of the above alternative repesentation to liao2020estimating. In view of the Bayesian framework, if we instead impose a prior $\pi_s$ that is supported on $A^{TI}\cap A^{EM-NTAT}$, the induced $\pi_F$ is supported on $\mathcal{F}$.

A Problem with the standard Bayesian method

Think of an econometrician who works with the $LATE$ framework and is aware of the testable implication of the IA-M assumption. She believes that there are some economic reasons that will lead to the presence of defiers. She also believes that the presence of defiers is abnormal and is unlikely to happen.

She decides to put a prior supported on the set $A^{TI}\cap A^{EM-NTAT}$, denoted by $\pi_{s}$. However, by doing so, she is forced to take a stance on other aspects of the distribution of potential outcomes. For example, the $\pi_{s}$ induces a distribution of the proportion of always takers: \[ \pi_{AT;s}(B)\equiv \int_{\mathcal{S}} \mathbbm{1}\left(E_{G^s}[\mathbbm{1}(D_{i}(1)=D_{i}(0)=1)]\in B \right)d\pi_s. \] The econometrician only believes that the presence of defiers is the cause of the possible model rejection, and she is worried that taking a stance on other aspects of the distribution of potential outcomes may lead to misleading or overconfident identified quantity. Therefore, the standard Bayesian method may not suit her goal in this context.

The robust Bayesian prior set

If a single prior can lead to over-specification of the prior belief, then the solution is to consider many prior beliefs to provide additional robustness, which is called the robust Bayesian method. We follow the robust Bayesian literature giacomini2021robust to assume that the observed data distribution $F$ is a reduced-form parameter, and the distribution $G^s$ is the structural parameter. Two structures $s$ and $s'$ are observationally equivalent if $M(G^s)=M(G^{s'})$. For any prior belief $\pi_s$, since $\mathcal{G}$ is a polish space with respect to the total variation metric\footnote{The polishness of the space $\mathcal{G}$ is sufficient for a well-defined conditional distribution $\pi_{s|F}$. See Chapter 2 of ghosal2017fundamentals.}, we can decompose the prior as:

equation[equation omitted — 93 chars of source]

where $\pi_F$ is a belief over the possible distributions of observed data, and $\pi_{s|F}$ is the conditional distribution of structures given observed $F$. For the belief to be consistent with the structural model, we require the posterior $\pi_{s|F}$ to be supported on the set $\{s: F=M(G^s)\}$. Following giacomini2021robust, we call the $\pi_{F}$ the updatable part because the belief of data distribution can be updated by observing the data realization. The $\pi_{s|F}$ is called the non-updatable part because given the $F$, data observation is ancillary to the conditional belief.

We assume that the econometrician has a unique belief marginal $\pi_F$ but can have multiple $\pi_{s|F}$ to be robust against the over-specification of the belief of structures. The uniqueness of $\pi_F$ does not harm the Bayesian modeling but helps with computation: The frequentist-Bayesian equivalence will ensure the posterior of $F$ does not depend on $\pi_F$ in the limit. Since the econometrician believes that the presence of defiers is the main reason for model rejection, she may only want to discipline the `No Defiers' aspect of the $G^s$ distribution. Given a prior $\pi_s$, let the conditional marginal distribution of amount of defiers be:

eqnarray[eqnarray omitted — 252 chars of source]

The econometrician forms a prior belief set $\Pi_s$ that has the following representation:

equation[equation omitted — 211 chars of source]

where $\pi_{DF|F}^*$ is the econometrician's choice of the belief of proportion of defiers. For example, if the econometrician believes that defiers are very unlikely to exists, then she can choose an exponentially decaying density function for $\pi_{DF|F}^*$. We use a single conditional marginal belief $\pi_{DF|F}^*$ in (ref), but the econometrician may experiment with multiple beliefs to gain additional robustness, which will be discussed in Section (ref).

The posterior distribution and interval estimates

The econometrician observes the realized data $\bm{X}_n=\{(Y_i,D_i,Z_i)\}_{i=1}^n$ and uses the Bayes' rule to update her belief about the real data distribution to be $\pi_{F|\bm{X}_n}$. Because of the decomposition (ref), the set of posterior is given by

equation[equation omitted — 196 chars of source]

The restrictions in the posterior set ((ref)) are similar to those in the prior set ((ref)) except the belief of the data distribution is updated to $\pi_{F|\bm{X}_n}$.

We propose the posterior upper- and lower-bound of the $LATE$ quantity:

equation[equation omitted — 295 chars of source]

and the following target of the confidence set:

equation[equation omitted — 160 chars of source]

which is a uniform confidence region for all possible posterior beliefs. For any value $\widehat{LATE}\in [LATE_*,LATE^*]$, we can find a posterior ${\pi_{s}\in \Pi_{s|\bm{X}_n}}$ such that $\widehat{LATE}$ is the posterior mean corresponding to $\pi_s$.

Characterization of the posterior quantities

For a given $\pi_s$, there is no closed-form expression for the $LATE^*$, $LATE_*$ and $CI$. The computation method in giacomini2021robust is developed for a finite-dimensional model and it can be hard to implement because the posterior $\pi_{s|\bm{X}_n}$ is a distribution over infinite dimensional objects. In this application, we use a combintation of analytical results and simulation to derive a feasible calculation of the bound $[LATE_*,LATE^*]$.

For any $\pi_{s|\bm{X}_n}\in \Pi_{s|\bm{X}_n}$, it can be written as the $\pi_{s|\bm{X}_n}=\pi_{s|F} \times \pi_{F|\bm{X}_n}$ for the uniquely chosen $\pi_{F|\bm{X}_n}$. Given the posterior for the data distribution $\pi_{F|\bm{X}_n}$, the variation in the posterior set $\Pi_{s|\bm{X}_n}$ comes from the variation in $\pi_{s|F}$, which is the `non-updatable' part. There are existing nonparametric Bayesian methods for us to find $\pi_{F|\bm{X}_n}$ and we can simulate $F$ from $\pi_{F|\bm{X}_n}$. We then use the potential outcome framework to analytically characterize the conditional distribution $\pi_{s|F}$ that achieves the upper and lower bound. The following Lemma states the feasibility of separating the simulation and analytic analysis parts.

propSuppose $-\infty<LATE_*\le LATE^*<\infty$, then the following equalities hold: \begin{equation} \begin{split} LATE^*=\int_{\mathcal{F}}\int_{\mathbb{R}^+} \overline{LATE}(F,m) d\pi^*_{DF|F}(m) d\pi_{F|\bm{X}_n},\\ LATE_*=\int_{\mathcal{F}}\int_{\mathbb{R}^+} LATE(F,m) d\pi^*_{DF|F}(m) d\pi_{F|\bm{X}_n}, \end{split} \end{equation} where \begin{equation} \begin{split} \overline{LATE}(F,m)&=\sup_{s\in A': m^{df}(G^s)=m, M(G^s)=F} E_{G^s}[Y_i(1)-Y_i(0)|D_i(1)=1,D_i(0)=0],\\ LATE(F,m)&=\inf_{s\in A': m^{df}(G^s)=m, M(G^s)=F} E_{G^s}[Y_i(1)-Y_i(0)|D_i(1)=1,D_i(0)=0], \end{split} \end{equation} where $A'=A^{TI}\cap A^{EM-ATNT}$.

Proposition (ref) separates the calculation of the $LATE$ bound into simulation part (ref) and the bound (ref). The bound (ref) requires us to find the structure $s$ that maximizes/minimizes the $LATE$ quantity given a fixed number of defiers and a fixed observed data disribution. In other words, for each possible value of the defier amount $m$, we find the most/least favorable structure $s$ and characterize the $LATE$ under this $s$.

We briefly discuss the intuition behind Proposition (ref). First, since we use a single $\pi_F$ and a single $\pi_{DF|F}^*$ to construct $\Pi_s$, by furthur conditioning on the value $m$, we have the decomposition $\pi_{s|\bm{X}_n}=\pi_{s|m,F}\times \pi^*_{DF|F}(m)\times \pi_{F|\bm{X}_n}$, which allows us to integrate over $\pi_{F|\bm{X}_n}$ and $\pi^*_{DF|F}$ in separate steps. It remains to optimize over $\pi_{s|m,F}$, which is equivalent to the pointwise optimization for each value of $m$ and $F$ as in (ref) .

It remains to characterize the optimization over $\{s: m^{df}(G^s)=m, M(G^s)=F\}$, which is hard to find since it is still an infinite dimensional optimization problem. We take a final step to transform the problem into a tractable finite-dimensional optimization problem.

thmLet $p_F(y,d)$ and $q_F(y,d)$ be the densities of $P_F(\cdot,d)$ and $Q_F(\cdot,d)$ with respect to the dominating measure $\mu_F$. We can compute $\overline{LATE}(F,m)$ using a 2-variable optimization program: \begin{equation} \overline{LATE}(F,m)=\sup_{a,b}\left[ \frac{\int_{\mathcal{Y}} y h_1^{max}(y)dy}{a} - \frac{\int_{\mathcal{Y}} yh_0^{max}(y)dy}{b} \right] \end{equation} subject to \begin{equation} \begin{split} &\quad aPr_F(Z_i=0)+Pr_F(D_i=1)-Pr_F(D_i=1|Z_i=1)Pr_F(Z_i=0)\\ &+bPr_F(Z_i=1)+Pr_F(D_i=0)-Pr_F(D_i=0|Z_i=0)Pr_F(Z_i=1)=m;\\ &\int_{\mathcal{Y}} \max\{p_F(y,1)-q_F(y,1),0\}dy \le a\le \int_{\mathcal{Y}} p_F(y,1) dy;\\ &\int_{\mathcal{Y}} \max\{q_F(y,0)-p_F(y,0),0\}dy \le b \le \int_{\mathcal{Y}} q_F(y,0)dy. \end{split} \end{equation} where $h_1^{max}$ and $h_0^{max}$ are specified in the following: \begin{equation} \begin{split} h_1^{max}(y)=\max\{p_F(y,1)-q_F(y,1),0\} + \min\{p_F(y,1),q_F(y,1)\}\mathbbm{1}(y> \bar{y}_{max})\\ h_0^{max}(y)= \max\{q_F(y,0)-p_F(y,0),0\}+ \min\{q_F(y,0),p_F(y,0)\}\mathbbm{1}(y< y_{max}) \end{split} \end{equation} with $\bar{y}_{max}$ and $\underline{y}_{max}$ be the solution to \[ \begin{split} \int_{\bar{y}_{max}}^{\infty} \min\{p_F(y,1),q_F(y,1)\} dy= a- \int_{\mathcal{Y}} \max\{p_F(y,1)-q_F(y,1),0\}dy\\ \int^{\underline{y}_{max}}_{-\infty} \min\{q_F(y,0),p_F(y,0)\} dy= b- \int_{\mathcal{Y}} \max\{q_F(y,0)-p_F(y,0),0\}dy. \end{split} \] Similarly, \[ \underline{LATE}(F,m)=\inf_{a,b}\left[ \frac{\int_{\mathcal{Y}} y h_1^{min}(y)dy}{a} - \frac{\int_{\mathcal{Y}} yh_0^{min}(y)dy}{b} \right] \] subject to the same linear constraint ((ref)). The $h_1^{max}$ and $h_0^{max}$ are specified in the following: \[ \begin{split} h_1^{min}(y)=\max\{p_F(y,1)-q_F(y,1),0\} + \min\{p_F(y,1),q_F(y,1)\}\mathbbm{1}(y< \bar{y}_{min})\\ h_0^{min}(y)= \max\{q_F(y,0)-p_F(y,0),0\}+ \min\{q_F(y,0),p_F(y,0)\}\mathbbm{1}(y> \underline{y}_{min}) \end{split} \] with $\bar{y}_{min}$ and $\underline{y}_{min}$ be the solution to \[ \begin{split} \int^{\bar{y}_{min}}_{-\infty} \min\{p_F(y,1),q_F(y,1)\} dy= a- \int_{\mathcal{Y}} \max\{p_F(y,1)-q_F(y,1),0\}dy\\ \int_{\underline{y}_{min}}^{\infty} \min\{q_F(y,0),p_F(y,0)\} dy= b- \int_{\mathcal{Y}} \max\{q_F(y,0)-p_F(y,0),0\}dy. \end{split} \]

The optimization problem (ref)-(ref) optimize over $a$ and $b$ and the constraints (ref) are in linear form can be implimented using standard optimization packages.

We provide some intuition for Theorem (ref). Let $h_1(y)$ denotes the density of $Y_i(1)$ and compliers conditional on $Z_i=1$,\footnote{That is, $h_1(y)=d Pr(Y_i(1)\le y,D_i(1)=1,D_i(0)=0|Z_i=1)/d\mu_F$.} and let $h_0(y)$ denotes the density of $Y_i(0)$ and compliers conditional $Z_i=0$. By definition of $LATE$, for this $h_1(y)$ and $h_0(y)$, we have $$LATE=\frac{\int_{\mathcal{Y}} yh_1(y)dy}{ \int_{\mathcal{Y}} h_1(y)dy} -\frac{ \int_{\mathcal{Y}} yh_0(y)dy}{ \int_{\mathcal{Y}} h_0(y)dy}.$$ At the same time, the amount of compliers, conditional on $Z_i=1$, is $\int_{\mathcal{Y}} h_1(y)dy$. So essentially, the optimization (ref) is converted to a two step optimization: 1.Given the value of compliers amounts for different $Z_i$ values (i.e., $a$ and $b$), what is the optimal densities; 2. What is the optimal $a$ and $b$ value.

The first question is answered by the $h_1^{max}(y)$ and $h_0^{min}(y)$. Fixing the $a$, to maximize $\int_{\mathcal{Y}} yh_1(y)dy/a$, we must allocate more probability masses to larger values of $y$. However, we also face the testable implication (ref) in the density form, and there is a minimum amount of density ($\max\{p_F(y,1)-q_F(y,1),0\}$) we must allocate to a $y$ value. In view of ((ref)), the first additive term of $h_1^{max}$ reflects this minimal level. After maintaining this essential density of compliers, we are left with $a- \int_{\mathcal{Y}} \max\{p_F(y,1)-q_F(y,1),0\}dy$ amount of compliers, which should be allocated to all values above the $\bar{y}_{max}$. At the same time, we cannot allocate too much density to large value of $h_1^{max}$, otherwise the testable implication for always takers will kick in, which is reflected in the second additive term of $h_1^{max}$ in ((ref)).\footnote{The maximal density for compliers conditional $Z_i=1$ is $p_F(y,1)$, and the difference between $\max\{p_F(y,1)-q_F(y,1),0\}$ and $p_F(y,1)$ is $\min\{p_F(y,1),q_F(y,1)\}$, which is the room for us to manipulate.} The rationale for $h_0^{max}$ is the same and we allocate all masses to small values of $y$. Given the value of $a$ and $b$, we can calculate $h_{1}^{max}$ and $h_{0}^{max}$ numerically.

The optimization of $a$ and $b$ does not have a closed-form solution and we have to use numerical optimization package. The first constraint in (ref) reflects that we have a fixed amount of defiers $m$, and the constraint is constructed via the potential outcome equation (ref). The second and third constraints come from the testable implication (ref).

Combining Proposition (ref) and Theorem (ref), we have a way to calculate the $LATE^*$, $LATE_*$ and the confidence set. We first generate a sample of $F$ from $\pi_{F|\bm{X}_n}$, and for each realized $F$, we generate a $m$ value from $\pi_{DF|F}^*$. Given the pair $(F,m)$, use Theorem (ref) to generate a sample of $\overline{LATE}(F,m)$ and $\underline{LATE}(F,m)$. The $LATE^*$ and $LATE_*$ are the simulated sample mean of $\overline{LATE}(F,m)$ and $\underline{LATE}(F,m)$ correspondingly. The confidence interval can also be construct from the distribution of $\overline{LATE}(F,m)$ and $\underline{LATE}(F,m)$ under $\pi_{F|\bm{X}_n}\times \pi_{DF|F}^*$. Let $\underline{C}_{\alpha/2}$ be the $\alpha/2$ quantile of $\underline{LATE}(F,m)$ under $\pi_{F|\bm{X}_n}\times \pi_{DF|F}^*$, and let $\bar{C}_{1-\alpha/2}$ be the $1-\alpha/2$ quantile of $\overline{LATE}(F,m)$ under $\pi_{F|\bm{X}_n}\times \pi_{DF|F}^*$. We use $[\underline{C}_{\alpha/2}, \bar{C}_{1-\alpha/2}]$ as the confidence interval.

propThe interval $CI\equiv[\underline{C}_{\alpha/2}, \bar{C}_{1-\alpha/2}]$ is a valid confidence interval in regard to ((ref)).

In Proposition (ref), we do not use the $\alpha$-quantile because we want to get some robustness when the upper and lower bound are very close to each other. However, when the posterior mean bounds are wide, the confidence interval can be conservative.\footnote{Additional moment-selection based andrews2010inference method can be adapted in the Bayesian inference context, but we leave it to keep the topic concentrated.}

Frequentist Equivalence

We now derive an equivalence result of our method to a frequentist approach to the refutable models. More equivalence results can be found in Section (ref). In the Bayesian statistics literature, the Berstein-von Mises theorem characterizes the equivalence between the Bayesian posterior distribution and the limit distribution under the frequentist world. In other words, Bayesian inference is equivalent to frequentist inference.

We do not seek to characterize the equivalence in the spirit of the Berstein-Von Mises theorem. This is because the structural model here is an infinite-dimensional and the nonparametric version of the Berstein-von Mises theorem is only valid under very specific models. Instead, we aim to show that the $LATE$ estimator under different frequentist models can be justified by a corresponding prior belief. As a result, we show that the frequentist assumption set is equivalent to a Bayesian prior belief specification.

We start with the crucial assumption on nonparametric posterior convergence result, and briefly describe the frequentists' minimal relaxation method.

assumptionLet $d_\mathcal{F}$ be the total variation metric on the $\mathcal{F}$. The posterior $\pi_{F|\bm{X}_n}$ converges weakly to the degenerate distribution with a point mass at $F_0$ along the empirical distribution sequence $\mathbb{P}_{F_0}^{(n)}$.
definitionLet $m_{min}^{df}(F)=\min_{s\in A^{TI}\cap A^{EM-NTAT}: F= M(G^s)} m^{df}(s)$ be the minimal probability of defiers that is required to rationalize $F$. We say a prior belief set $\Pi_s$ on $A^{TI}\cap A^{EM-NTAT}$ satisfies the minimal defier constraints if $\pi_{DF|F}^*= \delta(m_{min}^{df}(F))$, where $\delta(x)$ denotes the point mass measure at $x$.
propSuppose Assumption (ref) holds and $\mathcal{Y}$ is bounded. Let $\Pi_s^{Min}$ denote the prior set that satisfy the minimal defier constraints. Let $(Y_i,D_i,Z_i)\sim F_0$ be the true distribution of data, then the robust Bayesian estimator satisfies: \begin{equation} LATE^*=LATE_*\rightarrow_p \frac{\int_{\mathcal{Y}_1(F_0)} y(p_{F_0}(y,1)-q_{F_0}(y,1)) dy}{\int_{\mathcal{Y}_1(F_0)} (p_{F_0}(y,1)-q_{F_0}(y,1)dy} - \frac{\int_{\mathcal{Y}_0(F_0)} y(q_{F_0}(y,0)-p_{F_0}(y,0)) dy}{\int_{\mathcal{Y}_0(F_0)} (q_{F_0}(y,0)-p_{F_0}(y,0)dy}, \end{equation} where $\mathcal{Y}_1(F_0)=\{y\in\mathcal{Y}: p_{F_0}(y,1)\ge q_{F_0}(y,1)\}$ and $\mathcal{Y}_0(F_0)=\{y\in\mathcal{Y}: q_{F_0}(y,0)\ge p_{F_0}(y,0)\}$.

The limit in (ref) is the frequentists' minimal deviation method identified $LATE$ liao2020estimating,dahl2023nevertoolate. Proposition (ref) shows that the interval $ [LATE_*, LATE^*]$ shrinks to a point and converges to the minimal-deviation characterized identified quantity (ref). The result shows that the minimal-deviation method is equivalent to imposing a degenerate belief on the deviation from the `No Defiers' assumption. We use the Dirichlet process to model the prior distribution because it is known to have a consistenty result and computationally feasible posterior. We delegate the background of the Dirichlet process to ghosal2017fundamentals for readers who are not familiar with the nonparametric Bayes method, especially Chapters 2-4.

Posterior via Dirichlet Mixture Process Prior

We now propose a Bayesian estimation and inference procedure based on simulation. First, we specify a Dirichlet mixture process for the prior belief. The Dirichlet mixture process can accommodate density-based estimation than the Dirichlet process alone.

\paragraph{The prior of $F$} Weconsider a hierarchical prior to separate the discrete and the continuous part of $F$:

eqnarray[eqnarray omitted — 501 chars of source]

We briefly discuss the notation in ((ref))-((ref)) by comparing the notations to the parametric Bayesian method. We specify the marginal prior belief of $(D_i,Z_i)$ to satisfy a Dirichlet distribution with scale parameter $\alpha_{dz}$ and centering distribution $H_{dz}$. Here, $H_{dz}$ is a discrete measure supported on $\{0,1\}\times \{0,1\}$, which serves as the center of the Dirichlet distribution. The number $4$ specifies that $H_{dz}$ is supported on 4 discrete points. The finite measure $H_{dz}$ does not have to be a probability measure, but its total mass $|H_{dz}|\equiv H_{dz}(D_i\in\{0,1\},Z_i\in\{0,1\})$ is related to `variance' of the random draw. Roughly speaking, if $|H_{dz}|$ is larger, the prior belief weights more in the posterior. Any distribution drawn from $\text{Dirichlet} (4,H_{dz})$ can be viewed as a perturbation of the distribution $H_{dz}/|H_{dz}|$.

We use a different notation $\text{Dir}(\cdot)$ to denote the Dirichlet process. For any finite number of partitions of $\mathcal{Y}$, denoted by $\mathcal{Y}_1,...,\mathcal{Y}_K$, the finite discrete distribution generated from the Dirichlet process follows a Dirichlet distribution: $(\pi_{\theta|dz}(\mathcal{Y}_k)))_{k=1}^K\sim \text{Dirichlet}(K,(H_{\theta|dz}(\mathcal{Y}_k))_{k=1}^K)$. In other words, the Dirichlet process is the Kolmogrove-extension of the Dirichlet distribution. We specify the conditional prior belief of $Y_i|D_i,Z_i$ distribution as a Dirichlet process mixture as in ((ref)), which can be viewed as the weighted sum of densities for at different $\theta$ locations. Here $\psi_\sigma(y,\theta)$ is a kernel density function with auxiliary parameter $\sigma$. The parameter $\theta$ is assumed to follow a further Dirichlet process $\text{Dir} (H_{\theta|dz})$. The $Gamma(\alpha,\beta)$ distribution is chosen to be independent of the $Dir(H_{\theta|dz})$ process.\footnote{The additional parameter $\sigma$ serves as the mixture of Dirichlet Process Mixture. The additional mixture facilitates computation in many applications.} Depending on the application, $\psi$ can be chosen to be a normal pdf with mean $\theta$ and standard deviation $\sigma$ if $y$ is univariate and unbounded; When $y$ is bounded, $\psi_\sigma(y;\theta)$ can be chosen to be $Beta(a,b)$ function, where we choose $\alpha=0$\footnote{The choice of $\alpha=0$ is specific to the Beta kernel function and cannot be used for the Gaussian kernel.} so that $\sigma\equiv 0$ is a degenerate parameter, and $a<1\le b$ is sampled from $\theta=(a,b)\sim\pi_{\theta|dz}$.

\paragraph{Data-updated posterior} Recall that $\bm{X}_n=\{(Y_i,D_i,Z_i)\}_{i=1}^n$ denotes the observed data and let $\mathbb{F}^{DZ}_n$ denote the empirical distribution of data. The posterior $\pi_{dz|\bm{X}_n}$ of the distribution of $(D_i,Z_i)$ follows a Dirichlet distribution with parameter $(4, \frac{|H_{dz}|}{|H_{dz}|+n}H_{dz}+ \frac{n}{|H_{dz}|+n}\mathbb{F}^{DZ}_n)$. To calculate the conditional posterior $\pi_{y|D_i=d,Z_i=z,\bm{X}_n}$, we need to get the posterior of $\text{Dir} (H_{\theta|dz}|\bm{X}_n)$ and the posterior of $\Gamma(\alpha,\beta|\bm{X}_n)$. For the calculation of $\pi_{y|D_i=d,Z_i=z,\bm{X}_n}$, we will only use the observation with $D_i=d$ and $Z_i=z$. However, the conditional posterior $\pi_{y|D_i=d,Z_i=z,\bm{X}_n}$ does not have a closed-form solution but it can be calculated using the MCMC method, and an existing R package dirichletprocess can be used directly.

We will focus on the normal kernel $\psi$. The joint posterior $\pi_{F|\bm{X}_n}$ is known to satisfy the consistency Assumption (ref) under mild conditions.

assumptionThe true densities $p_{F_0}(y,d)$ and $q_{F_0}(y,d)$ satisfy the following entropy constraints for $d\in\{0,1\}$:\begin{enumerate} • $\int_{\mathcal{Y}} p_{F_0}(y,d) \log p_{F_0}(y,d) dy<\infty$ and $\int_{\mathcal{Y}} q_{F_0}(y,d) \log q_{F_0}(y,d) dy<\infty$; • $-\int_{\mathcal{Y}} p_{F_0}(y,d) \log\left[\inf_{||y'||<\delta}p_{F_0}(y-y',d) \right] dy<\infty $, and \\$-\int_{\mathcal{Y}} q_{F_0}(y,d) \log\left[\inf_{||y'||<\delta}q_{F_0}(y-y',d) \right] dy<\infty $ for some $\delta>0$. \end{enumerate} Moreover, the parameter $\alpha>1$ in the Gamma distribution.
propAssumption (ref) implies that the posterior estiamted under ((ref))-((ref)) satisfies Assumption (ref) for the normal kernel $\psi_{\sigma}(\cdot;\theta)$.

Proposition (ref) establish the posterior consistency result and also validate the frequentist equivalence result in Proposition (ref). We conclude the section by summarizing the poterior sampling altorithm.

algorithm[algorithm omitted — 1,729 chars of source]

Empirical Application

In this section, we apply the robust Bayesian method to card1993using, who studied the causal effect of college attendance on earnings. We take the outcome variable $Y_i$ to be an individual $i$'s log wage in 1976, and $D_i=1$ to be individual $i$'s four-year college attendance. We use the college proximity variable $Z_i$ as an instrument for college attendance, i.e. $Z_i=1$ means the individual was born near a four-year college. The empirical setting has been used by both kitagawa2015 and mourifie2017testing to test the IA-M assumption. In various settings, the IA-M assumption is rejected, and it is reasonable to believe that defiers are likely to present.

We follow mourifie2017testing to condition $(Y_i,D_i,Z_i)$ on three characteristics: living in the south (S/NS), living in a metropolitan area (M/NM), and an African-American ethnic group (B/NB). We also drop the NS/NM/B and NS/M/B group due to a small sample size or a small number of $Z_i=0$.

For the choice of prior parameters in ((ref))-((ref)), we choose $H_{dz}$ to be the uniform distribution over $\{0,1\}\times \{0,1\}$, and $\psi(\cdot;\theta)$ to be normal density function with $\theta=(\mu,\sigma)$. The centering measure $H_{(\mu,\sigma)|dz}$ is assumed to follow the standard bivariate normal distribution. We choose the parameter $(\alpha,\beta)=(2,4)$ which is the default value in the R package. We choose $\pi_{DF|F}^*$ to be a Gaussian-decaying density $\pi_{DF|F}^*(m)=\frac{C}{\sqrt{2\pi}\sigma(F)}e^{-(m-m^{min}(F))^2/\sigma^2(F)}$ on $[m^{min}(F),m^{max}(F)]$, zero otherwise, where $m^{min}(F)$ and $m^{max}(F)$ are correspondingly the minimal and maximal amount of defiers required to justify the $F$ distribution\footnote{That is $m^{min}(F)=\min\{m^{df}(s): F=M(G^s)\}$. Similar definition holds for $m^{max}(F)$.}. We choose $\sigma(F)=(m^{max}(F)-m^{min}(F))/1.96$. The constant $C$ is chosen to make $\pi_{DF|F}^*$ a proper density function. The estimation and inference results are shown in Table (ref). {

table[table omitted — 1,643 chars of source]

}

There are several interesting observations from Table (ref). First, the robust Bayesian interval estimate is more stable than the $LATE^{wald}$. The $LATE^{wald}$ can estimate very negative $LATE$ (S/NM/NB and S/M/NB groups) or unrealistic high $LATE$ (S/M/B). In comparison, the set estimator under the robust Bayesian method is more persistent across different groups. We also compare the result under the robust Bayesian method to that under the frequentist minimal defiers, see Proposition (ref). The set estimator and the confidence set under the robust Bayesian method are wide. Therefore, frequentists' identified set under the minimal-defier Assumption (see Proposition (ref)) may generate over-precise $LATE$ estimates. The wider $LATE$ bound in the robust Bayesian framework shows that the quantity is very sensitive to the presence of defiers, and the frequentist minimal defiers approach can generate over-confident interpretation of the return of college education.

A General Theory

In this section, we develop a robust Bayesian theory for models with refutable assumptions that generalizes the insight from the $LATE$ application to incomplete models. We continue using the notations in the $LATE$ model and revisit the $LATE$ application whenever a new concept is introduced. We delegate further illustrating examples to Section (ref).

Definitions

Starting from a vector of observed variables $X$, we can define the observation space $\mathcal{F}$ as the collection of all regular distributions of $F(X)$. The regularity condition, such as the smoothness of $F(X)$ or existence of certain moments of $X$, are conviently chosen by the econometrician to analyze the model.\footnote{These conditions, such as exitence of moments or densities, are often non-testable given a finite sample.}

Following koopmans1950identification and jovanovic1989, we consider that any outcome $X$ is generated through some distribution of underlying random vector $\epsilon$ through some mapping $M$. A pair of a distribution of $\epsilon$, denoted by $G$, and a mapping $M$ is called an econometric structure. The distribution $G$ governs the fundamental heterogeneity at the observation level. The mapping $M$ governs the economic machenism that generates the observed data $X$ from $\epsilon$. Since in most econometric problems, we focus on the distribution of outcomes $F$ instead of how each $X$ is related to $\epsilon$, we directly define the mapping $M$ as a mapping from the space of underlying variable distributions to $\mathcal{F}$.

definitionAn econometric structure $s=(G^s,M^s)$ consists of a distribution $G^s$ of $\epsilon$, and an outcome mapping $M^s$. Let $\mathcal{G}$ denote the space of all regular distributions of $G^s(\epsilon)$. The outcome mapping $M^s$ is a function $M^s:\mathcal{G}\rightarrow \mathcal{F}$.

The requirement of the regularity of $G^s$ is in the same spirit as that of $F$. Unlike in the $LATE$ example where $M$ is fixed for all strutures, the outcome mapping $M^s$ can also vary across econometric structures to capture differences in economic mechanisms. Think of an OLS regression $Y_i=\beta^s_0+\beta^s_1 X_i +\eta_i$. A structure consists of a distribution $G^s$ of $(X_i,\eta_i)$ and a mapping $M^s$ which can be characterized by $(\beta^s_0,\beta^s_1)$.

definitionA structure universe $\mathcal{S}$ is a collection of structures such that $\cup_{s\in\mathcal{S}} M^s(G^s)=\mathcal{F}$, and an assumption $A$ is a subset of $\mathcal{S}$.

The structure universe determines the paradigm that can encompass different contexts while $A$ imposes assumptions that are convenient for a particular empirical context. In view of the OLS regression where we believe the average effects of $X_i$ on $Y_i$ is positive, the $\mathcal{S}$ does not restrict the parameter space, but $A=\{s: \beta_1^s\ge 0\}$ is a restriction. An assumption $A$ is called refutable if there exists some observed distribution $F\in\mathcal{F}$ such that $F\notin\cup_{s\in A} M^s(G^s)$. In other words, we can find an observed data distribution that cannot be generated by econometric structures inside $A$. In this case, we are aware of the possible violation of $A$, and a prior belief that putting all probability on $A$ is inappropriate. Instead, we may consider that $A$ is violated, but the violation of $A$ is a rare event and correspondingly form a prior belief.

To formalize the construction of the prior belief set, we first consider that the original assumption can be written as the intersection of countably many sub-assumptions $A=\cap_{j=1}^\infty A_j$, and for each assumption, we find a deviation metric $m_j(s):\mathcal{S}\rightarrow \mathbb{R}_+$ that measures the deviation of structure $s$ from $A_j$. Moreover, we require that the deviation metric satisfies \[ A=\{s: m_j(s)=0\}\cap (\cap_{l\ne j} A_{l} ). \] This representation above requires that $m_j(s)$ is a sharp characterization of $A_j$ given the rest of the assumptions, which allows us to use $m_j(s)$ to describe our belief of deviation from the baseline model $A$. The choice of $A_j$, $m_j$ and the multiplicity of ways to relaxed assumption are discussed in liao2020estimating.\footnote{{Also see Appendix (ref) for an additional example that illustrates the multiplicity issue.}}

Robust Bayesian Prior

The econometrician wants to take a stance on assumption $A_j$ while leaving holding all other assumptions untouched. To do so, she puts a prior $\pi_{s}$ over the rest of the assumptions $\cap_{l\ne j} A_{l}$. The prior belief $\pi_{s}$ induces a marginal distribution on $\mathcal{F}$ via \[ \pi_F (\mathcal{F}_0)=\int_{\mathcal{S}} \mathbbm{1}(M^s(G^s)\in \mathcal{F}_0) d\pi_{s}. \] Unlike finite dimensional models where metrics on econometric structures are natrual, when we extend the parametric analysis to nonparametric models, we need to further make topological assumptions on $\mathcal{S}$ and $\mathcal{F}$. We endow the structural space with a metric $d_s$ and the observed distribution space $\mathcal{F}$ with a metric $d_\mathcal{F}$. When $(\mathcal{S},d_s)$ and $(\mathcal{F},d_F)$ are polish spaces, conditional distribution is well-defined\footnote{See Chapter 2 of ghosal2017fundamentals.}, and the prior belief $\pi_s$ can be characterized as the product of two measures \[ \pi_s= \pi_{s|F}\times \pi_{F}. \] Since we want to maintain the rest of the assumptions, we assume that $\pi_s$ is supported on $\cap_{l\ne j} A_{l}$. The econometrician may not want to put a single $\pi_s$ because by doing so she also takes stances on the likelihood of other aspects of the model.\footnote{In the $LATE$ example, a single prior also specifies the distribution of always takers. In the OLS regression, putting a single prior not only specifies the likelihood of deviation from $\{s: \beta_1^s\ge0\}$ but also the distribution of $\beta_0^s$. } Instead, she believes that only $A_j$ is likely to be violated and is only willing to make a statement with respect to the distribution of $m_j(s)$. In this case, we propose a robust Bayesian prior set:

equation[equation omitted — 180 chars of source]

where $\pi_{m_j|F}$ is the marginal distribution of the deviation from $A_j$ conditional on the $F$, i.e. $\pi_{m_j|F}(B)\equiv \int_{\mathcal{S}} \mathbbm{1}(m_j(s)\in B) d \pi_{s|F} $ for any measurable set $B$. The general prior set ((ref)) is slightly different from (ref): We allow the marginal belief of deviation from $A_j$ to fall in a general set $\Pi_{m_j|F}$ rather than choosing a single marginal belief. The set $\Pi_{m_j|F}$ specifies the econometricians' belief of the deviation from the $A_j$ assumption and allows for additional robustness against the choice of prior. The set of beliefs can change with $F$ to allow for flexibility of specification of beliefs. The following are some examples of the possible choices of $\Pi_{m_j|F}$:

equation[equation omitted — 433 chars of source]

where $a^F_{min}$ and $a^F_{max}$ are the $F$ induced minimal and maximal deviations.\footnote{In the LATE example, the constraints on $a$ and $b$ in (ref) set the bounds for probability of defiers.} The first set is the collection of all possible prior beliefs that have a decreasing density of deviation, this is a large set and may impose too few restrictions for the model to generate informative results. The second one is a uniform prior while the third is a truncated normal density. We can also use convex combination of two sets to create a new prior set: $\Pi^{mix}_{m_j|F}=\alpha\Pi_{m_j|F}^2+(1-\alpha)\Pi_{m_j|F}^3$, for $\alpha\in[0,1]$.

Characterization of the posterior of parameters

The econometrician is interested in an 1-dimensional parameter $\theta$, which is a function from the structure space to the parameter space $\theta(s): \mathcal{S} \rightarrow \Theta \subseteq \mathbb{R}$.

After observing data realization $\bm{X}_n$, the econometrician's posterior belief set is given by \[ \Pi_{s|\bm{X}_n}=\{\pi_{s|\bm{X}_n}:\pi_{s|\bm{X}_n}=\pi_{s|F}\times \pi_{F|\bm{X}_n}, \quad \pi_{s|\bm{X}_n}\text{ supported on } \cap_{l\ne j} A_{l},\quad \pi_{m_j|F}\in \Pi_{m_j|F}\} \] The above posterior set shows the observed data only updates the belief of data distribution $F$, and it leaves the conditioning belief $\pi_{s|F}$ unchanged. The posterior beliefs set then induce a set of beliefs of the parameter of interest:

equation[equation omitted — 259 chars of source]

The corresponding bound estimator of the parameter of interest is given by \[ \theta^*=\sup_{\pi_{\theta|\bm{X}_n}\in \Pi_{\theta|\bm{X}_n}}E_{\pi_{\theta|\bm{X}_n}}[\theta],\quad \theta_*=\inf_{\pi_{\theta|\bm{X}_n}\in \Pi_{\theta|\bm{X}_n}}E_{\pi_{\theta|\bm{X}_n}}[\theta]. \] For inference, we seek a confidence set $CI$ such that \[ \inf_{\pi_{\theta|\bm{X}_n}\in \Pi_{\theta|\bm{X}_n}} Pr_{\pi_{\theta|\bm{X}_n}} (\theta\in CI)\ge 1-\alpha. \] The optimization with respect to the $\pi_{\theta|\bm{X}_n}$ is an infinite dimensional optimization problem and can be hard to characterize. Instead, we can transform the optimization problem to optimize the parameter value given a particular deviation value. Compared to the posterior set (ref) in the $LATE$ example, the general posterior set (ref) may contains multiple conditional marginal distribution of deviation $\pi_{m_j|F}$. We first discuss the case where $\Pi_{m_j|F}$ is a singleton and then think about more general cases.

propFor a given deviation value $m_j(s)=m$, and the observed data distribution $F$, define the intermediate quantity \[ \begin{split} \bar{\theta}(F,m)=\sup_{s: m_j(s)=m,M^s(G^s)=F} \theta(s),\\ \underline{\theta}(F,m)=\inf_{s: m_j(s)=m,M^s(G^s)=F} \theta(s). \end{split} \] Suppose $\Pi_{m_j|F}$ conatins a singleton $\pi_{m_j|F}^*$, $-\infty<\theta_*\le \theta^*<\infty$, and $\bar{\theta}(F,m), \underline{\theta}(F,m)$ are integrable with respect to $\pi^*_{m_j|F}\pi_{F|\bm{X}_n}$, then the $\theta^*$ and $\theta_*$ can be characterized by \[ \begin{split} \theta^*= \int_{\mathcal{F}}\int_{\mathbb{R}^+}\bar{\theta}(F,m) d\pi^*_{m_j|F} d\pi_{F|\bm{X}_n},\\ \theta_*= \int_{\mathcal{F}}\int_{\mathbb{R}^+} \underline{\theta}(F,m) d\pi^*_{m_j|F} d\pi_{F|\bm{X}_n}.\\ \end{split} \]

Proposition (ref) breaks the optimization problem into model analysis part ($ \bar{\theta}(F,m),\underline{\theta}(F,m)$) and the integration problem. The integration problem can be solved by simulating $(m,F)$ from the product distribution $\pi_{m_j,s|F}^*\times\pi_{F|\bm{X}_n}$. The model analysis part does not have a general solution and is model specific. In the $LATE$ example and subsequently in Section 4, we show that even for nonparametric models, characterizing $\bar{\theta}(F,m)$ and $\underline{\theta}(F,m)$ is feasible. The feasibility comes from the fact that we are not giving up $A_j$ completely: The assumption $A_j$ is initially imposed to facilitate identification of $\theta$ and often a closed-form identified expression is available. When we deviate from $A_j$ by $m_j(s)=m$, such a deviation often results in a tractable change of the econometric structure from the closed-form $A_j$-identified quantity and it is possible to reduce the nonparametric optimization problem to a finite dimensional optimization problem.\footnote{For the $LATE$ application, this is shown in Theorem (ref). For an application to the discrete choice model that relaxes the Logit error, see Proposition (ref).}

propSuppose the conditions in Proposition (ref) hold. Let $C_{\alpha/2}$ be the lower $\alpha/2$ quantile of the distribution of $\underline{\theta}(F,m)$, and let $C_{1-\alpha/2}$ be the upper $\alpha/2$ quantile of the distribution of $\bar{\theta}(F,m)$, with respect to the product probability $\pi_{m_j|F}\times \pi_{F|\bm{X}_n}$. Then $[C_{\alpha/2},C_{1-\alpha/2}]$ is a valid confidence interval.

We now consider the case when $\Pi_{m_j|F}$ contains mutiple prior beliefs. If $\Pi_{m_j|F}$ contains finite many beliefs, then we can simply repeat Proposition (ref) finitely times and take the union bound. Note that we only need to calculate $\bar{\theta}(F,m)$ and $\underline{\theta}(F,m)$ once and they can be applied for different conditional marginal distribution of deviation. In the following proposition, we show that it is possible to characterize the bounds for a convex prior set.

propSuppose $\Pi_{m_j|F}^s$ is a convex set, then in Proposition (ref), we just need to focus on the extreme points of $\Pi_{m_j|F}^s$, i.e., we have: \[\theta^*=\sup_{\pi_{\theta|\bm{X}_n}\in extrm\left(\Pi_{\theta|\bm{X}_n}\right)}E_{\pi_{\theta|\bm{X}_n}}[\theta],\quad \theta_*=\inf_{\pi_{\theta|\bm{X}_n}\in extrm\left(\Pi_{\theta|\bm{X}_n}\right)}E_{\pi_{\theta|\bm{X}_n}}[\theta],\] where $extrm(\Pi_{\theta|\bm{X}_n})$ is the set of extreme points of $\Pi_{\theta|\bm{X}_n}$.

Proposition (ref) provides a way to compute $\theta^*$ and $\theta_*$ for many complex prior sets. For example, the set of decreasing conditional marginal density of deviation $\Pi_{m_j|F}^1$ in (ref) has the extreme point sets characterized by step functions.\footnote{Consider the set of weakly decreasing functions defined on $[a,b]$ whose integral is 1, and equip this set with the $||\cdot||_\infty$ norm. Then the extreme points of this set are step functions $f(x)=\mathbbm{1}(a\le x\le a+c)/c$. }

Equivalence to Frequentists' Approaches

We now extend the propositions in Section (ref) to the general framework. We start with the general nonparametric Bayesian convergence assumption. Equip $\mathcal{F}$ with a metric $d_\mathcal{F}$. Let $F_0$ be the true data distribution such that $X_i\sim F_0$, and let $\delta_{F_0}$ be the degenerate distribution with point mass at $F_0$. We use the following Bayesian posterior convergence definition, see Proposition 6.2 in ghosal2017fundamentals. Throughout this section, we maintain that the true econometric structure $s_0$ (in the frequentists' framework) satisfies $s_0\in\cap_{l\ne j} A_{l}$.

assumptionThe posterior distribution $\pi_{F|\bm{X}_n}$ converges weakly to the $\delta_{F_0}$ along the sequence of $\mathbb{P}_{F_0}^{(n)}$, where the $\mathbb{P}_{F_0}^{(n)}$ is the sampling probability measure.\footnote{It is the probability measure of the $\{X_i\}_{i=1}^n$.}

Assumption (ref) is the high-level convergence assumption on the posterior. Recall that the definition weak convergence depends on the topology of the sample space $\mathcal{F}$, and hence $d_\mathcal{F}$ matters: If we only care about the CDF, then $d_\mathcal{F}$ can be the Kolmogorov-Smirnov distance; If we care about the density estimation, then $d_\mathcal{F}$ can be chosen as the total variation distance (or equivalently the $L_1$-distance of the density of $F$).

Equivalence to the minimal deviation method.

We start with the definition of the minimal deviation method and the prior set that is equivalent to this minimal deviation method.

definitionLet $m_{j,min}(F)=\min_{s: F=M^s(G^s)} m_j(s)$ be the minimal amount of deviations that is required to rationalize the data distribution $F$. We say a prior $\pi_s$ on $\cap_{l\ne j} A_{l}$ satisfies the minimal deviation constraints if the induced conditional marginal distribution $\pi_{m_j|F}$ has a point mass on $m_{j,min}(F)$ for all $F$.
assumptionThe bounds $\bar{\theta}(F,m_{j,min}(F))$ and $\underline{\theta}(F,m_{j,min}(F))$ are continuous at $F_0$.

Assumption (ref) is a high-level assumption but it should be easy to pin down by more fundamental conditions. For example, if we can show that $ \bar{\theta}(F,m)$ is a continuous function, and $m_{j,min}(F)$ is a continuous function, then Assumption (ref) follows for $\bar{\theta}(F,m_{j,min}(F))$.

propLet the true data distribution be $F_0$ and the conditions in Proposition (ref) hold. Define the minimal-deviation identified set $\Theta^{ID,min}(F_0)=\{\theta(s): F_0\in M^s(G^s), m_j(s)=m_{j,min}(F_0) \}$. Suppose the identified set $\Theta^{ID,min}(F_0)$ is bounded, then for any prior $\pi_s$ that satisfies the minimal deviation constraint and Assumption (ref), we can find some large value $M>0$ such that \[ d_H([\max\{\theta_*,-M\},\min\{\theta^*,M\}], conv\left(\Theta^{ID,min}(F_0)\right)) \rightarrow_p 0, \] where $d_H$ is the Hausdorff metric, $conv\left(\Theta^{ID,min}(F)\right)$ is the convex hull of the identified set, and $\rightarrow_p$ is convergence in probability with respect to the frequentist sampling probability $\mathbb{P}_{F_0}^{(n)}$.

The trimming number $M$ in Proposition (ref) helps to avoid irregular tail behaviors in the nonparametric posterior set $\Pi_{\theta|\bm{X}_n}$. The intuition of Proposition (ref) is simple: If we put a point mass prior at the minimal deviation amount that rationalize the data, there is no variation in the deviation amount and it is equivalent to let the data to select the minimal deviation amount $(m_{j,min}(F_0))$ and identify $\theta$. However, there are two additional features in Proposition (ref). First, Assumption (ref) requires the bounds to be continuous as a function of $F$. Intuitively, if the bounds are not continuous, then the weak convergence of the posterior distribution to $F_0$ does not imply the convergence in mean.\footnote{The continuity condition is also crucial for frequentists' estimator to be consistent.} Second, the robust Bayesian posterior bound converges to the convex hull of the frequentists' identified set, but the convexification cannot be avoided because the posterior expectation can be viewed as the Aumann expectation of the identified set giacomini2021robust.

Equivalence to giving up $A_j$

Another frequentists' approach is to completely give up $A_j$ and identify the parameter of interest under $\cap_{l\ne j} A_{l}$. We show that the frequentist's identified set under $\cap_{l\ne j} A_{l}$ can also be rationalized by a robust prior set.

definitionLet $m_{j,max}(F)=\min_{s: F=M^s(G^s)} m_j(s)$ be the maximal amount of deviation that can be used to rationalize the data distribution $F$. We say a prior set $\Pi_{m_j|F}$ satisfies the giving up $A_j$ constraint if it is the class of all distributions supported on $[m_{j,min}(F),m_{j,max}(F)]$.

In the definition of $\Pi_{m_j|F}$ above, we simply give up all restrictions on the conditional marginal distribution of deviation except the support condition that is required to ensure the prior set $\Pi_{m_j|F}$ is proper.\footnote{By definition, it is impossible to find a prior belief $\pi_s$ supported on $\cap_{l\ne j} A_{l}$ whose conditional distribution $\pi_{m_j|F}$ with a support wider than $[m_{j,min}(F),m_{j,max}(F)]$. }

assumptionThe support bounds $m_{j,min}(F)$ and $m_{j,max}(F)$ are continuous in $F$. Moreover, the function $\theta^*$. Moreover, $ \bar{\theta}(F,m)$ and $\underline{\theta}(F,m)$ are continuous in $(m,F)$.
propLet $\Theta^{ID}_{-j}(F)=\{\theta(s): s\in \cap_{l\ne j} A_{l}, F=M^s(G^s)\}$ be the frequentist's identified set under $\cap_{l\ne j} A_{l}$. Let $\Pi_{m_j|F}$ be the prior set that satisfies the giving up $A_j$ constraint. If $ \sup \Theta^{ID}_{-j}(F)$ and $ \inf \Theta^{ID}_{-j}(F)$ are continuous in $F$, then there exists a large value of $M>0$ such that \[ d_H([\max\{\theta_*,-M\},\min\{\theta^*,M\}], conv\left(\Theta^{ID}_{-j}(F_0)\right)) \rightarrow_p 0. \]

In addition to the continuous bounds assumption, we also require the support bounds to be continuous in $F$. This is because we want to ensure the prior set does not change drastically when we change $F$ slightly.

Additional Applications

We now examine an application to intersection bounds models and an application to discrete choice models with Logit errors. We illustrate the usefulness of the robust Bayesian method in dealing with refutable models. Appendix (ref) uses the application to monotone IV models manski1998monotone to illustrate subtle issues in choosing the deviation metric $m_j$.

In view of Proposition (ref), we aim to show that $\bar{\theta}(F,m)$ and $\underline{\theta}(F,m)$ can be characterized by closed-form expression or by feasible computational method.

Intersection Bounds and Moment Inequality Models

We start with a simple intersection bounds model chernozhukov2013intersection. Suppose we have access to observed variables including bound variables $(\bar{Y}_i,\underline{Y}_i)$ and an instrument $Z_i\in [z_l,z_u]$. The observed variables can be rationalized by the following model: $\bar{Y}_i=Y_i+\eta_i^+$, $\underline{Y}_i=Y_i+\eta_i^-$, where the underlying variables $\epsilon_i=(Y_i,\eta_i^+,\eta_i^-,Z_i)$. We are interested in the mean of $Y_i$, that is $\theta(G)=E_{G}[Y_i]$. The rationale behind the intersection bound model is to assume that the instrument $Z_i$ does not the mean of $Y_i$, but it may influence the bounds via $\eta_i^+,\eta_i^-$.

assumptionThe intersection bound assumption $A=A_1\cap A_2$, where $A_1$: $E_G[Y_i|Z_i]=E_G[Y_i]$; and $A_2$: $\inf_{z\in [z_l,z_u]} E_G[\eta_i^+|Z_i=z]\ge0\ge \sup_{z\in [z_l,z_u]} E_G[\eta_i^-|Z_i=z]$.

Under the intersection bound assumption, we can derive that, for any $F$ and $F=M^s(G^s)$, we must have $\theta(G)\in [\sup_z E[\underline{Y}_i|Z_i=z], \inf_z E[\bar{Y}_i|Z_i=z]]$. The model is refuted if and only if the interval is empty. We focus on relaxing the $A_2$ assumption while maintaining $A_1$. A simple way to measure the deviation from $A_2$ is the following metric: \[ m_2(G)= \left( \sup_{z\in [z_l,z_u]} E_G[\eta_i^-|Z_i=z]\right)_+ +\left(-\inf_{z\in [z_l,z_u]} E_G[\eta_i^+|Z_i=z]\right)_+, \] where $(t)_+=\max\{0,t\}$. We derive the bounds of $\theta(G)$ in the view of Proposition (ref).

propFor any observed data distribution $F$, the minimal deviation required to rationalize $F$ is given by $m^{min}_2(F)\equiv \left(\sup_{z\in [z_l,z_u]} E_F[\underline{Y}_i|Z_i=z]-\inf_{z\in [z_l,z_u]} E_F[\bar{Y}_i|Z_i=z]\right)_+$. For any $m\ge m^{min}_2(F)$, we have \[ \begin{split} &\underline{\theta}(F,m)= \sup_{z\in [z_l,z_u]} E_F[\underline{Y}_i|Z_i=z]-m,\\ &\bar{\theta}(F,m)= \inf_{z\in [z_l,z_u]} E_F[\overline{Y}_i|Z_i=z]+m. \end{split} \]

The intersection bounds model implies a moment inequality model where we can write the model as $E[\bar{Y}_i-\theta|Z_i]\ge 0$ and $E[\theta-\underline{Y}_i|Z_i]\ge 0$. A `reduced-form' way to relax the assumption is to assume $E[\bar{Y}_i-\theta|Z_i]\ge -m$ while keeping $E[\theta-\underline{Y}_i|Z_i]\ge 0$. In this case, we can characterize $ \bar{\theta}(F,m)= \inf_z E_F[\bar{Y}_i|Z_i=z]+m$ and $\underline{\theta}(F,m)= \sup_z E[\underline{Y}_i|Z_i=z]$. Though relaxing the moment inequality model results in the same bounds on $\theta(G)$ for some choice of deviation metric,\footnote{ Relaxing $E[\bar{Y}_i-\theta|Z_i]\ge -m$ is equivalent to consider a deviation metric $m_j(G)=\inf_{z} \left[E[-\eta_i^+|Z_i=z]\right]_+$.} we do not recommend directly relaxing the moment inequality bound: Recall that in Section (ref), we define $m_j$ as a function on structures rather than on the observed distribution $F$, because we want to ensure the economic interpretation of the deviation value $m$. In contrast, a distance metric defined on $\mathcal{F}$ may not have a direct interpretation from the econometric structures.

Discrete Choice and Relaxing Distributional Assumptions

We now apply our method to study the distributional relaxation in christensen2023counterfactual. We consider the discrete choice model where we observe: (1).$Y_i\in \{0,1,2,...,J\}$, the discrete choice outcome; and (2).$Z_i=\{0,1\}$, an indicator of whether choice $J$ is available for individual $i$, and $Z_i=1$ means choice $J$ is available for individual $i$. The standard discrete choice model assumes individual $i$ has a random utility of choosing $j$: $U_{ij}=u_j +\xi_{ij}$, where $\bm{u}\equiv(u_1,...,u_J)$ is the mean utility that does not change across individuals. Choice $j=0$ is the outside option whose utility is normalized to zero. Consumes choose the object that maximize her utility.

In this example, observed variables are $X_i=(Y_i,Z_i)$ whose distribution is $F$, and underlying variables are $(\xi_{ij},Z_i)$ whose distribution is $G$. A structure $s$ consists of $s=(\bm{u}^s,G^s)$.\footnote{The mapping $M^s$ is determined by $Y_i=\arg\max_j {u_j+\xi_{ij}}$, and $M^s$ can be summarized by the mean utility vector $\bm{u}$. } We are interested in the mean utility parameter for the last choice ($\theta(s)=u_J^s$). We follow christensen2023counterfactual to impose the following moment conditions for econometric structure $s$:

equation[equation omitted — 267 chars of source]

The $P_{j|z}$ is an additional nuisance parameter introduced in christensen2023counterfactual, which is identified as $E_F[\mathbbm{1}(Y_i=j)|Z_i=z]$ from the data distribution $F$. In the robust Bayesian framework, $P_{j|z}$ will be estimated by Bayesian methods.

An empirically convenient assumption is to assume that the random utility shocks $\xi_{ij}$ are i.i.d drawn from the Type-I extreme value, which delivers a closed-form expression of the choice probability. However, the Type-I extreme value assumption also induces the I.I.A property on the data identified $P_{j|z}$. Since we have the variation in choice set due to the variation in $Z_i$, the model can be refuted if the I.I.A property fails.

christensen2023counterfactual propose to relax the Type-I extreme value by consider the set of distributions: \[ \mathcal{G}_m=\left\{G: D(G||G_0)\le m\right\}, \] where $G_0$ is the Type-I extreme value distribution, $D(G||G_0)$ is the chi-squared divergence between $G$ and $G_0$.\footnote{The chi-squared divergence $D(G||G_0)=\frac{1}{2}\left[\int \frac{(dG)^2}{d(G_0)}d\mu -1\right]$, where $dG$ is the Radon-Nikodym derivatives of $G$ with respect to the dominating measure $\mu$. We do not use the Kullback-Leibler divergence because it is not well defined for Type-I extreme value distribution. See Section 6 of christensen2023counterfactual for a discussion.} Our goal is to characterize the parameter bound at a particular Chi-squared divergence level $m$, i.e., we focus on the boundary set $\partial \mathcal{G}_m=\left\{G: D(G||G_0)= m\right\}$ and we want to characterize the bounds:

equation[equation omitted — 379 chars of source]

Optimization problems ((ref)) are infinite dimensional, but we can use the duality in christensen2023counterfactual to recast the problem into a feasible finite dimensional optimization.

propSolving optimization problems ((ref)) are equivalent to solving the \begin{equation} \begin{split} \theta(F,m)=\inf u^s_J\quad s.t.\quad m\in (\Delta(\bm{u}^s,\bm{P}),\bar{\Delta}(\bm{u}^s,\bm{P})),\\ \bar{\theta}(F,m)=\sup u^s_J\quad s.t.\quad m\in (\Delta(\bm{u}^s,\bm{P}),\bar{\Delta}(\bm{u}^s,\bm{P})). \end{split} \end{equation} where $\bm{P}$ is the vector corresponds to the $P_{j|z}$ in ((ref)) and depends on $F$ implicitly, and \begin{equation} \begin{split} \Delta(\bm{u}^s,\bm{P})&=\sup_{\zeta\in \mathbb{R},\lambda\in \mathbb{R}^{2J-1}} -E_{G_0}\left[\phi^\star\left(-\zeta-\lambda'\bm{g}(\bm{u}^s,\bm{\xi})\right)\right]-\zeta-\lambda'\bm{P} \quad s.t.\quad ((ref)) holds,\\ \bar{\Delta}(\bm{u}^s,\bm{P})&=\inf_{\zeta\in \mathbb{R},\lambda\in \mathbb{R}^{2J-1}} E_{G_0}\left[\phi^\star\left(\zeta+\lambda'\bm{g}(\bm{u}^s,\bm{\xi})\right)\right]+\zeta+\lambda'\bm{P} \quad \quad \,\,\, s.t.\quad ((ref)) holds, \end{split} \end{equation} where $\phi^\star(x)=0.5x^2+x$ is the dual function for the Chi-squared divergence, $\bm{g}(\bm{u}^s,\bm{\xi})$ is the vectorized moment function in ((ref)), whose dimension\footnote{Note that the moment equality for $j=0$ is redundant since the choice probabilities sum to 1. Also, we have one less choice for the $z=0$ case.} is $2J-1$, and $E_{G_0}$ takes the expectation of $\bm{\xi}$ under the standard i.i.d Logit distribution. In addition, if the maximum or minimum can be achieved in ((ref)), the corresponding open intervals in ((ref)) should be changed to closed intervals.\footnote{For example, if $ \underline{\Delta}(\bm{u}^s,\bm{P})$ can be achieved by some $\zeta,\lambda$ value while $\bar{\Delta}(\bm{u}^s,\bm{P})$ cannot be achieved, then the intervals in ((ref)) should be changed to $[\underline{\Delta}(\bm{u}^s,\bm{P}),\bar{\Delta}(\bm{u}^s,\bm{P}))$. }

It's notable that both ((ref)) and ((ref)) are finite dimensional optimization problems. Given a fixed $\bm{u}^s$, christensen2023counterfactual interpret $\underline{\Delta}(\bm{u}^s,\bm{P})$ as the minimal Chi-squared divergence relaxation of the $G_0$ that is required to rationalize $\bm{u}^s$ with moment condition ((ref)), and $\bar{\Delta}(\bm{u}^s,\bm{P})$ is the corresponding maximal Chi-squared divergence relaxation. The optimization (ref) says that as long as the deviation $m$ is between the minimal and maximal deviation for $\bm{u}^s$, then $\bm{u}^s$ can be rationalized by some underlying distribution $G^s$ with $G^s\in \partial \mathcal{G}$. We then subsequently optimize over $u_J$ to find the bound. Proposition (ref) is not a trivial corollary of christensen2023counterfactual, and (ref) is the additional optimization that we need to compute in the robust Bayesian framework.

Conclusion

This paper proposes a robust Bayesian approach to deal with econometric models with refutable assumptions. The robust Bayesian approach considers a set of prior beliefs that share the same conditional marginal distribution of deviation from the refutable assumption. We propose combining a simulation method and model analysis methods to compute the posterior mean set and the confidence set for the parameter of interest.

As leading applications of the robust Bayesian methods, we study the $LATE$ model, the intersection bounds model, and the discrete choice model with Logit errors. In all three models, it is possible to reduce the infinite-dimensional problems of model analysis to finite-dimensional computational problems. The nice property is likely because of the closed-form characterization of the identified parameter under the refutable models.

There are several problems that we do not discuss and leave for possible future work. First, we focus on complete models where $M^s$ is a function. In many economic models, such as discrete game models, multiple outcomes can arise, which may require us to define $M^s$ as a multi-valued correspondence. Second, it is interesting to think about the Berstein-von Mise results for the parameter of interests. Note that even if we start with an infinite dimensional model, the parameter of interest is a finite-dimensional parameter, whose asymptotic property may be derived with stronger statistical assumptions. Third, extension to multi-dimensional parameter of interests can be done at the cost of more complicated computation of the $\bar{\theta}(F,m)$ and $\underline{\theta}(F,m)$.