EconBase
← Back to paper

Causal Inference under Outcome-Based Sampling with Monotonicity Assumptions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

71,683 characters · 13 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

-1cm Causal inference under outcome-based sampling with monotonicity assumptions

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 104 chars of source]

} \fi

abstractWe study causal inference under case-control and case-population sampling. Specifically, we focus on the binary-outcome and binary-treatment case, where the parameters of interest are causal relative and attributable risks defined via the potential outcome framework. It is shown that strong ignorability is not always as powerful as it is under random sampling and that certain monotonicity assumptions yield comparable results in terms of sharp identified intervals. Specifically, the usual odds ratio is shown to be a sharp identified upper bound on causal relative risk under the monotone treatment response and monotone treatment selection assumptions. We offer algorithms for inference on the causal parameters that are aggregated over the true population distribution of the covariates. We show the usefulness of our approach by studying three empirical examples: the benefit of attending private school for entering a prestigious university in Pakistan; the relationship between staying in school and getting involved with drug-trafficking gangs in Brazil; and the link between physicians’ hours and size of the group practice in the United States.

{\it Keywords:} Relative risk; attributable risk; odds ratio; partial identification

\spacingset{1.8}

Introduction

Random sampling is convenient for causal inference, but it may be too costly in practice for various reasons. For instance, rare events are likely to be severely under-represented in a random sample of a finite size: e.g., cancer breslow1980statistical, infant death currie2005air, consumer bankruptcy Domowitz:1999, entering a highly prestigious university Delavande:Zafar:19, and drug trafficking carvalho2016living. The objective of this paper is to study causal inference in outcome-based sampling scenarios such as case-control or case-population studies.

We focus on observational data, as opposed to experimental ones, with a binary outcome and a binary treatment. Holland:Rubin adopt the potential outcome framework to show that the assumption of strong ignorability can be used to identify the counterfactual odds ratio in case-control studies. They then argue that the counterfactual odds ratio approximates the ratio of two potential-outcome probabilities (i.e., causal relative risk) under the rare-disease assumption, which says that the probability of outcome occurrence (e.g., having “a certain disease”) is close to zero. Their work is our starting point, and we make additional contributions in several ways.

First, we focus on two direct causal parameters (i.e., a ratio or a difference of two potential-outcome probabilities) that are more straightforward to interpret than counterfactual odds ratios: our parameters of interest are the causal relative and attributable risks given a specific value of covariates. Second, we do not appeal to the rare-disease assumption, and we take the perspective of partial identification manski1995book,manski2003partial,manski2009identification,tamer2010partial,molinari:2020. Third, we consider a set of monotonicity assumptions, and we compare their identification power with that of strong ignorability. Strong ignorability is a popular setup for causal inference, but its identification power in outcome-based sampling turns out to be somewhat limited. Specifically, in case-control or case-population studies, strong ignorability is generally not sufficient to point identify the causal relative and attributable risks. We can obtain bounds on them, but they are not much better than those we can obtain in a less restrictive setup using monotonicity. Specifically, we will consider monotone treatment response manski1997 and monotone treatment selection manski2000monotone.

Our work builds upon manski2009identification, who conducts a partial identification analysis for both relative and attributable risks under outcome-based sampling without focusing on causal parameters. The MTR and MTS assumptions as well as other related notions of monotonicity have been extensively used in the literature. For example, see vytlacil2007dummy, bhattacharya2008treatment,BSV-2012, pearl2009causality, vanderweele2009propertise, Kreider-et-al:2012, jiang2014monotone, okumura2014concave, choi2017estimation, Kim-et-al:2018, and MSV-2019 among others.

We now discuss the relation of our work with the existing literature on causal inference under outcome-based sampling. maansson2007estimation point out that the propensity score method has only limited ability to control for confounding factors in case-control studies. Our method does not rely on the propensity score. rose2011causal and van2011targeted use an assumption that the true case probability is known by a prior study. We focus on the instance of unknown case probability. CC-handbook-chapter provide an extensive survey on causal inference in case-control studies, but no discussion on partial identification approaches can be found there. Therefore, possibilities based on partial identification appear to be rather underexplored. kuroki2010sharp and gabriel2020causal are notable exceptions. gabriel2020causal obtain bounds on the causal attributable risk in a variety of scenarios including outcome-dependent sampling with an instrumental variable. But they do not leverage any monotonicity assumption, while we do not consider instrumental variables but we use monotonicity restrictions. kuroki2010sharp is more similar to our work in that they obtain bounds on both the causal relative and attributable risks by using the MTR assumption. Our contributions relative to kuroki2010sharp can be highlighted as follows: (1) we exploit not only the MTR but also the MTS assumption, and therefore the bounds are different; (2) we consider case-control sampling as well as case-population sampling; (3) we compare the identification power of the popular assumption of strong ignorability with that of the MTR and MTS assumptions; (4) we consider how to aggregate the causal parameters over the distribution of the covariates; and (5) we provide algorithms for causal inference.

The remaining part of the paper is organized as follows. In (ref) we formally present the setup including the causal parameters of interest and the sampling schemes. (ref) focus on the causal relative risk and attributable risk to address identification and aggregation. (ref) covers how to carry out causal inference. (ref) presents three empirical applications. Specifically, by using datasets collected in previous studies, we address new research questions that are not examined in the original papers. All the proofs, discussions on semiparametric efficiency and computational algorithms are in Online Appendix. An accompanying R package is available on the Comprehensive R Archive Network (CRAN) at \href{https://CRAN.R-project.org/package=ciccr}{https://CRAN.R-project.org/package=ciccr}, and the replication files are available at \href{https://github.com/sokbae/replication-JunLee-JBES}{https://github.com/sokbae/replication-JunLee-JBES}.

Preliminaries

Causal parameters

Let $(Y^*, T^*, X^*)$ be a random vector of a binary outcome, a binary treatment, and covariates of a representative individual. Since we are interested in outcome-based sampling, we assume that a random sample of $(Y^*, T^*, X^*)$ is not available. Instead, we have a sample of $(Y,T,X)$, where the distribution of $(T,X)$ given $Y$ is related with that of $(T^*,X^*)$ given $Y^*$. The exact sampling schemes and related assumptions will be discussed in detail later, and in this subsection, we only focus on the parameters of interest.

For the sake of causal inference, we use the usual potential outcome notation. So, $Y^*(t)$ will be the potential outcome for $T^*=t$, and $Y^*$ can be written as $Y^* = Y^*(1)T^* + Y^*(0)(1-T^*)$. Therefore, our notation extends chen2001parametric and Xie-et-al:JASA by adding an extra layer of potential outcomes. The causal effect of the treatment can be measured by either (conditional) relative risk or attributable risk: each of them is defined as follows:

align[align omitted — 247 chars of source]

provided that the denominator of $\theta_\mathrm{RR}(x)$ is strictly positive. Therefore, $\theta_\mathrm{AR}(x)$ is the usual conditional average treatment effect, whereas $\theta_\mathrm{RR}(x)$ is a causal version of the relative risk parameter.

Relative risk defined by a ratio of “success” probabilities has been popular in epidemiology and biostatistics, particularly when the “success” is a rare event: if the treatment changes the success probability from $0.01$ to $0.02$, then it is a 100% increase, though the difference of $0.01$ may suggest an impression that the change was unimportant. Further, it turns out that $\theta_\mathrm{RR}(x)$ is closely related with the odds ratio (in terms of the observed variables), which has been widely used as a measure of association in case-control studies.

Bernoulli sampling

As we mentioned earlier, we assume that a random sample of $(Y^*, T^*, X^*)$ is unavailable. Instead the researcher has access to a random sample of $(Y,T,X)$, where the distribution of $(Y,T,X)$ is related with that of $(Y^*,T^*,X^*)$ by Bernoulli sampling breslow2000semi that we describe below.

In Bernoulli sampling, the researcher first draws a Beroulli variable $Y$ from a pre-specified marginal distribution, after which she randomly draws $(T,X)$ from some $\mathcal{P}_y$ if and only if $Y=y$; so, $Y$ is an artificial device to decide which subpopulation we will draw $(T,X)$ from. If $\mathcal{P}_y$ is the distribution of $(T^*,X^*)$ conditional on $Y^*=y$, then this is nothing but case-control sampling. Since $h_0 = \mathbb{P}(Y=1) \in (0,1)$ is part of the sampling scheme, we will assume that it is known; if not, it can be easily estimated without compromising inferential validity. See online appendices B.1 and B.2 for more details. Before we proceed, we make a common support assumption for simplification.

assumption[Common Support] The support of $X^*$ and that of $X$ given $Y=y$ for $y=0,1$ coincide; the common support will be denoted by $\mathcal{X}$.

Below we discuss two leading cases of Bernoulli sampling that we focus on throughout the paper.

design[Case-Control Sampling] For $y\in\{0,1\}$, $\mathcal{P}_y$ is the distribution of $(T^*,X^*)$ given $Y^*=y$.
design[Case-Population Sampling] $\mathcal{P}_1$ is the conditional distribution of $(T^*,X^*)$ given $Y^*=1$, whereas $\mathcal{P}_0$ represents the distribution of $(T^*,X^*)$ of the entire population.

(ref) is arguably the most popular form of case-control studies breslow1996statistics and (ref) was referred to as “contaminated case-control studies” by lancaster1996case: we call the latter design case-population sampling, which is more descriptive. The case-population sampling design has been used to study drug trafficking carvalho2016living and mass demonstrations rosenfeld_2017 among others.

Note that the distribution of $(T,X)$ is identified from the data, but that of $(T^*,X^*)$ may not. For instance, in (ref), we have $f_X(x) = f_{X^*|Y^*}(x\mid 1)h_0 + f_{X^*|Y^*}(x\mid 0)(1-h_0) \neq f_{X^*}(x)$, unless $h_0$ is the same as $p_0:=\mathbb{P}(Y^*=1)$, i.e., the true probability of the case in the population. Further, $f_{YX}(1,x) = f_{X^*|Y^*}(x\mid 1) h_0 = f_{X^*}(x) \mathbb{P}(Y^*=1\mid X^* = x) h_0/p_0$, which yields the likelihood function studied in e.g.\ Manski:Lerman. We emphasize that $\mathbb{P}(Y=1\mid X=x)$ does not have economic interpretation like $\mathbb{P}(Y^* = 1\mid X^*=x)$, where the latter is often specified via domain knowledge in a specific field such as a utility function with an additively separable normal or Gumbel error term.

Identification

In this section, we study identification of the causal parameters, i.e., $\theta_\mathrm{RR}(x)$ and $\theta_\mathrm{AR}(x)$. Aggregation over $x\in\mathcal{X}$ will be considered later. For the purpose of the identification analysis, we consider two sets of assumptions: one is the standard case of strong ignorability, and the other is an alternative possibility based on monotonicity assumptions. We will see that even strong ignorability is not sufficient to point-identify $\theta_\mathrm{RR}(x)$ or $\theta_\mathrm{AR}(x)$ under case-control sampling, i.e., (ref).

We consider the following assumptions.

assumption[Overlap] For all $(y,t,s,x) \in \{0,1\}^3 \times \mathcal{X}$, \begin{equation*} 0 < \mathbb{P}\{ Y^*(t) = y, T^* = s \mid X^* = x\} < 1. \end{equation*}
assumption[Unconfoundedness] For all $(t,x)\in\{0,1\}\times \mathcal{X}$, \begin{equation*} \mathbb{P}\{ Y^*(t) = 1\mid T^*=1, X^*=x \} = \mathbb{P}\{ Y^*(t) = 1\mid T^*=0, X^*=x \}. \end{equation*}

(ref) together constitute strong ignorability, which is a standard setup for causal inference. (ref) is stated in terms of the joint probability mass function of $Y^*(t)$ and $T^*$ given $X^*=x$. We do this for a few reasons. First, (ref) ensures that all the conditional probabilities we consider and their ratios are well-defined: e.g., $\theta_\mathrm{RR}(x)$ is well-defined under (ref). Also, it ensures that the distribution of $(Y,T,X)$ has enough overlap to identify $\mathbb{P}(T=t\mid Y=y, X=x)$ under each of the two Bernoulli sampling schemes.

The key component of the strong ignorability setup is (ref). In the following subsections we will start from clarifying how far (ref) can take us to identify the causal parameters under case-control and case-population sampling. Although it is standard, strong ignorability does not allow the treatment assignment to be endogenous. Therefore, we consider a set of alternative assumptions under which we study how much we can say about the causal parameters under the two sampling scenarios.

assumption[Monotone Treatment Response] $Y^*(1)\geq Y^*(0)$ almost surely.
assumption[Monotone Treatment Selection] For all $t\in \{0,1\}$ and $x \in \mathcal{X}$, \begin{equation*} \mathbb{P}\{ Y^*(t) = 1\mid T^* = 1, X^* = x\} \geq \mathbb{P}\{ Y^*(t) = 1 \mid T^* = 0, X^* = x\}. \end{equation*}

(ref) was first proposed by manski1997, while (ref) was used by manski2000monotone. (ref) says that treatment is potentially beneficial but it never hurts. For instance, if an individual does not earn high income with a college degree, then the person will not be highly paid without a college degree, either. (ref) states that, all else being equal, individuals with a higher degree are at least as likely to earn high incomes if their educational attainment were randomly assigned, as compared to those without a higher degree. In essence, the treatment decision made by an individual reveals their `type': continuing with the same example, those opting for a higher degree are more motivated, and they would be at least as likely to earn high incomes as those who choose not to pursue a higher degree if they were randomly assigned to different educational attainment. (ref) is trivially weaker than (ref), and it allows individuals with `higher ability' to self-select a higher degree.

Before we move on, we define the following functions:

align[align omitted — 284 chars of source]

where $\mathbb{P}(Y=1\mid X=x)$ is the prospective regression function identified from the data. Here, both $r_\mathrm{CC}(x,p)$ and $r_\mathrm{CP}(x,p)$ can be alternatively expressed by using the conditional densities of $X$ given $Y=y$ by the Bayes rule, which is related with the distribution of $X^*$ given $Y^*=y$ or simply the distribution of $X^*$, depending on the sampling design. Indeed, it can be shown that $r_\mathrm{CC}(x,p_0) = \mathbb{P}(Y^*=1\mid X^*=x)$ under case-control sampling and $r_\mathrm{CP}(x,p_0) = \mathbb{P}(Y^*=1\mid X^*=x)$ under case-population sampling, where $p_0 = \mathbb{P}(Y^*=1)$: see lemma A.3 in the online appendix. Therefore, one can view the functions $r_\mathrm{CC}$ and $r_\mathrm{CP}$ as devices to exploit the fact that the only unidentified object in our context will be $p_0$.

Causal relative risk

In this section, we study identification of $\theta_\mathrm{RR}(x)$, for which we first introduce some notation. Let $\Pi(t\mid y,x) = \mathbb{P}(T=t \mid Y=y, X=x)$ be the retrospective regression function. For $(x,p)\in \mathcal{X}\times [0,1]$ and for $d\in \{\mathrm{CC},\mathrm{CP}\}$, define \[ \Gamma_{d,\mathrm{RR}}(x,p) := \frac{\Pi(1\mid 1,x)}{\Pi(0\mid 1,x)}\times \frac{ \Pi(0\mid 0,x) + r_d(x,p) \{\Pi(0\mid 1,x) - \Pi(0\mid 0, x)\} }{ \Pi(1\mid 0,x) + r_d(x,p) \{\Pi(1\mid 1,x) - \Pi(1\mid 0, x)\} }, \] where (ref) ensures that $\Pi(t\mid y,x) \neq 0$ for all $(t,y,x)\in \{0,1\}\times \{0,1\}\times \mathcal{X}$ in each of the two Bernoulli sampling schemes. It is worth noting that $\Gamma_{d,\mathrm{RR}}(x,0)$ for both $d\in\{\mathrm{CC},\mathrm{CP}\}$ is just the covariate-adjusted odds ratio, i.e., \[ \mathrm{OR}(x) := \frac{\Pi(1\mid 1,x)}{\Pi(0\mid 1,x)}\frac{\Pi(0\mid 0,x)}{\Pi(1\mid 0,x)} , \] which is a popular measure of covariate-adjusted association in case-control studies. Since $\mathrm{OR}(x)$ is more descriptive than $\Gamma_{d,\mathrm{RR}}(x,0)$, we will use the former notation whenever it is relevant.

The following lemma shows what we could achieve if we had a random sample, i.e., if $(Y^*,T^*,X^*)$ were observed.

lemma[RR-Benchmark] If (ref) are satisfied, then for all $x\in \mathcal{X}$, \[ \theta_\mathrm{RR}(x) = \frac{\mathbb{P}(Y^*=1\mid T^*=1, X^*=x)}{\mathbb{P}(Y^*=1\mid T^*=0, X^*=x)}. \] Alternatively, if (ref) are satisfied, then for all $x\in\mathcal{X}$, \[ 1\leq \theta_\mathrm{RR}(x) \leq \frac{\mathbb{P}(Y^*=1\mid T^*=1, X^*=x)}{\mathbb{P}(Y^*=1\mid T^*=0, X^*=x)}, \] where the bounds are sharp.

(ref) serves two purposes. First, it is useful as a middle step to establish the sharp identifiable bounds under Bernoulli sampling, i.e., under designs (ref) (Case-Control) and (ref) (Case-Population). Second, it shows benchmark results for the identification of $\theta_\mathrm{RR}(x)$ in that it shows the best we can achieve under random sampling through unconfoundedness or monotonicity. Therefore, (ref) should be compared with (ref) that are discussed below.

Point identification under random sampling and strong ignorability is not surprising. Partial identification under random sampling and the monotonicity assumptions is reminiscent of e.g.\ manski2000monotone. However, in our setup, the researcher does not have access to a random sample of $(Y^*,T^*,X^*)$, and therefore, (ref) is not an identification result. It will serve as a benchmark to show the cost of case-control or case-population studies in terms of identification.

Recall that $p_0 = \mathbb{P}(Y^*=1)$ is the true probability of the case, which is an unidentified object under Bernoulli sampling.

theoremSuppose that (ref) are satisfied. Then, for all $x\in \mathcal{X}$, we have the following. \begin{enumerate}[(1)] • Under case-control sampling, i.e., (ref), we have $\theta_\mathrm{RR}(x) = \Gamma_{\mathrm{CC},\mathrm{RR}}(x, p_0)$. • Under case-population sampling, i.e., (ref), we have $\theta_\mathrm{RR}(x) = \mathrm{OR}(x)$. \end{enumerate}

(ref) is not identification results in the case of case-control sampling: $p_0$ is unidentified in (ref). In contrast, it shows that $\theta_\mathrm{RR}(x)$ is point identified under case-population sampling. Therefore, (ref) provides an easier environment for causal inference, at least under unconfoundedness. It seems ironic that (ref) was referred to as case-control sampling with contamination by lancaster1996case but that the `contamination' is in fact helpful for identification.

In the case of (ref) we do not have point identification, but there is only one simple parameter that is unidentified. Therefore, it is not too difficult to proceed with a partial identification approach. We will further elaborate about this possibility. Before we proceed though, it is worth comparing the case-control case of (ref) with Holland:Rubin. Specifically, Holland:Rubin show that under (ref), $\mathrm{OR}(x)$ is equal to the odds ratio in terms of the potential outcomes if strong ignorability is imposed: i.e.,

equation[equation omitted — 229 chars of source]

(ref) is an identification result, but its right-hand side expression is not straightforward to interpret. It appears that the reason Holland:Rubin emphasized the right-hand side expression of (ref) instead of the more easily interpretable causal relative risk $\theta_\mathrm{RR}(x)$ is that the former is identified by $\mathrm{OR}(x)$, whereas the latter necessitates addressing the issue that $p_0$ remains unidentified.

Generally, $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,p_0)$ is different from $\mathrm{OR}(x) = \Gamma_{\mathrm{CC},\mathrm{RR}}(x,0)$. However, this issue has been traditionally ignored, because if $Y^*$ represents a rare event in that $p_0\approx 0$, then $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,p_0) \approx \Gamma_{\mathrm{CC},\mathrm{RR}}(x,0)$ by continuity: the assumption of small $p_0$ is known as the rare disease assumption in epidemiology. However, the quality of the approximation via continuity can quickly decrease as $p_0$ deviates from zero, i.e., the occurrence of $Y^*=1$ becomes less uncommon in the population. Therefore, when $p_0$ is away from zero, a natural alternative approach is to take a partial identification approach, where we target the function $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,\cdot)$ itself, at least within a certain neighborhood of $0$.

Below we will write $f_{A,B}(a,b)$ for the Radon-Nikodym density of $A,B$ (with respect to some dominating measure). For instance, when $A$ is discrete and $B$ is continuous, we will have $f_{A,B}(a,b) = \mathbb{P}(A=a) f_{B|A}(b\mid a)$ by using a product of count and Lebesgue measures. Similarly, $f_{A,B|C}(a,b\mid c)$ will be used to denote a conditional density of $(A,B)$ at $(a,b)$ given $C=c$.

assumptionThere is a known value $\bar p$ such that $p_0 \leq \bar p$, where $\bar p \leq 1$ under (ref), and $\bar p \leq \bar p^*$ with \begin{equation} \bar p^* := \inf\ \Biggl\{ \frac{f_{T,X|Y}(t,x\mid 0)}{f_{T,X|Y}(t,x\mid 1)}:\ $t,x$ are such that $f_{T,X|Y}(t,x\mid 1)>0$ \Biggr\} \end{equation} under (ref).

We remark that $\bar p^*\leq 1$: see online appendix F. Under case-control or case-population sampling, $p_0 = \mathbb{P}(Y^*=1)$ is generally unidentified, because $Y^*$ is not randomly observed. Since case-control or case-population sampling is popular when $Y^*=1$ is a rare event and therefore a random sample of a modest size tends to contain too few observations of the case of interest, we do not want to rule out the possibility that $p_0$ is close to zero: it is straightforward though to replace (ref) with the one that $p_0\in [\underaccent{\bar}{p}, \bar p]$ for some known values of $\underaccent{\bar}{p}$ and $\bar p$.

If we have an auxiliary sample, from which we learn about $p_0$, then plugging that piece of information into the case-control sample will resolve the identification problem since $p_0$ is the only unidentified object here. Even if it is difficult to pin down $p_0$ exactly, we may have external sources or qualitative information about how prevalent a certain `disease' is, and such information can be used to place an upper bound on $p_0$. Relying on the researcher's prior knowledge on an unidentified object has been used in the context of robust estimation as well horowitz1995identification,horowitz199716.

Choosing $\bar p = 1$ in (ref) corresponds to the case where the researcher has no prior information for $p_0$ at all: we do not rule out this possibility. In (ref), it may be possible to find $\bar p<1$ even without having any external source of information at all. To see this point, we note that under (ref), we must have

align*[align* omitted — 252 chars of source]

where $f_{T^*,X^*|Y^*}(t,x\mid 0)(1-p_0) = f_{T,X|Y}(t,x\mid 0) - f_{T,X|Y}(t,x\mid 1)p_0 \geq 0$ for all $t,x$. This motivates the definition of $\bar p^*$ in (ref).

theoremSuppose that (ref) are satisfied. Under case-control sampling, i.e., (ref), we have \[ \min\{ \mathrm{OR}(x),\ \Gamma_{\mathrm{CC},\mathrm{RR}}(x,\bar p)\} \leq \theta_\mathrm{RR}(x) \leq \max\{ \mathrm{OR}(x),\ \Gamma_{\mathrm{CC},\mathrm{RR}}(x,\bar p)\}, \] and the bounds are sharp.

(ref) is a simple corollary from (ref), where it is addressed that $p_0$ is unidentified under case-control sampling. Since $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,p)$ is monotonic in $p\in [ 0, \bar p ]$, it suffices to consider the two end points to obtain sharp bounds, where one of the end points is the odds ratio $\mathrm{OR}(x) = \Gamma_{\mathrm{CC},\mathrm{RR}}(x,0)$. We also remark that it can be verified that $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,\bar p) \geq 0$ because $0\leq r_\mathrm{CC}(x,\bar p)\leq 1$ by definition: this should not be surprising because $\theta_\mathrm{RR}(x) \geq 0$ by definition.

If (ref) is satisfied in addition, then we can show that $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,\cdot)$ is a decreasing function and therefore it follows that $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,\bar p) \leq \theta_\mathrm{RR}(x) \leq \Gamma_{\mathrm{CC},\mathrm{RR}}(x,0) = \mathrm{OR}(x)$ under (ref). Therefore, the odds ratio represents the maximum causal relative risk that is consistent with what is observed in a case-control study. If there is no information for $p_0$ at all, then the lower bound is simply one. Below we will see that the sharp identifiable bounds $[1,\ \mathrm{OR}(x)]$ on $\theta_\mathrm{RR}(x)$ can still be obtained without relying on the ignorability assumptions in case-control studies.

Unconfoundedness is a popular assumption for causal inference, but it is not always satisfied in observational studies. Further, unlike the standard case of random sampling, it does not deliver point-identification under case-control studies. (ref) provide an alternative possibility, where we do not lose much in terms of partial identification.

theoremSuppose that (ref) are satisfied. Then, under both (ref), we have $1\leq \theta_\mathrm{RR}(x) \leq \mathrm{OR}(x)$, where the bounds are sharp.

Unlike (ref), (ref) considers the case where we do not have unconfoundedness but we only impose monotonicity. Now, $\mathrm{OR}(x)$ is a sharp upper bound on $\theta_\mathrm{RR}(x)$ under both case-control and case-population sampling designs.

It is not explicit in (ref), but its proof shows that the knowledge of $p_0$ is potentially useful in (ref) but not in (ref). In fact, if $p_0$ were known, then the sharp bounds on $\theta_\mathrm{RR}(x)$ under (ref) would be given by $[ 1,\ \Gamma_{\mathrm{CC},\mathrm{RR}}(x,p_0) ]$, whereas those under (ref) would still be $[1,\ \mathrm{OR}(x)]$. This difference arises because a few applications of the Bayes rule show that the sharp upper bound under random sampling, i.e., the prospective regression ratio $\mathbb{P}(Y^*=1\mid T^*=1,X^*=x)/\mathbb{P}(Y^*=1\mid T^*=0,X^*=x)$ in (ref), is equal to $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,p_0)$ under (ref), whereas it is equal to $\Gamma_{\mathrm{CP},\mathrm{RR}}(x,0) = \mathrm{OR}(x)$ under (ref). Therefore, if we do not have a random sample, but we have access only to a case-control sample, then there is an information loss in terms of sharp identifiable bounds on $\theta_\mathrm{RR}(x)$. In contrast, a case-population sample is equally informative for $\theta_\mathrm{RR}(x)$ as a random sample. Thus, (ref) provides a better environment for causal inference than (ref) under monotonicity, similarly to the case of unconfoundedness: see our comments below (ref). The extra challenge in case-control studies can be addressed by the fact that $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,p)$ is decreasing in $p$. Therefore, the sharp upper bound on $\theta_\mathrm{RR}(x)$ under (ref) is given by the maximum (over $p$) of $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,p)$, which is equal to $\Gamma_{\mathrm{CC},\mathrm{RR}}(x,0)$ even without using (ref).

We now compare (ref) with (ref). The identification power of strong ignorability depends on the specific sampling design, whereas that of the monotonicity assumptions is independent of which of the two sampling scenarios applies. Specifically, in case-population studies, i.e., (ref), unconfoundedness is informative in that it ensures that $\theta_\mathrm{RR}(x)$ is point identified by the odds ratio. However, in case-control studies, i.e., (ref), unconfoundedness only yields interval identification, where the sharp identifiable bounds are the same as what the monotonicity assumptions can deliver if we have no information for $p_0$.

Causal attributable risk

We now turn to the alternative causal parameter $\theta_\mathrm{AR}(x)$. We need some extra notation. For $(x,p) \in \mathcal{X}\times [0,1]$ and for $d\in \{\mathrm{CC},\mathrm{CP}\}$, define \[ \Gamma_{d,\mathrm{AR}}(x,p) := \sum_{j=0}^1 \frac{(-1)^{j+1}\Pi(j\mid 1,x)}{\Pi(j\mid 0,x) + r_d(x,p)\{\Pi(j\mid 1,x) - \Pi(j\mid 0, x)\}}, \] where $\Pi(t\mid y,x)$ and $r_d(x,p)$ are defined in the beginning of (ref). Note that $\Gamma_{d,\mathrm{AR}}(x,0)$ is not exactly the odds difference, though it is similar: it is a difference between two ratios of retrospective regressions.

We start with the benchmark case of what if we could observe $(Y^*,T^*,X^*)$.

lemma[AR-Benchmark] If (ref) are satisfied, then for all $x\in \mathcal{X}$, \[ \theta_\mathrm{AR}(x) = \mathbb{P}(Y^*=1\mid T^*=1, X^*=x) - \mathbb{P}(Y^*=1\mid T^*=0, X^*=x). \] Alternatively, if (ref) are satisfied, then for all $x\in \mathcal{X}$, \[ 0\leq \theta_\mathrm{AR}(x) \leq \mathbb{P}(Y^*=1\mid T^*=1, X^*=x) - \mathbb{P}(Y^*=1\mid T^*=0, X^*=x), \] where the bounds are sharp.

Similarly to (ref), (ref) has two purposes. First, it is a middle-step result to establish the sharp identifiable bounds on $\theta_\mathrm{AR}(x)$ when we do not have a random sample but only a sample from either (ref) or (ref) is available. Second, it shows benchmark results for the identification of $\theta_\mathrm{AR}(x)$ via unconfoundedness or monotonicity under random sampling. Point identification of $\theta_\mathrm{AR}(x)$ via strong ignorability under random sampling is now a standard result. If strong ignorability is replaced with the monotonicity assumptions, then the regression difference should be interpreted as a sharp upper bound on the causal attributable risk. Below we extend these results to the cases of case-control and case-population sampling.

theoremSuppose that (ref) are satisfied. Then, for all $x\in \mathcal{X}$, we have the following. \begin{enumerate}[(1)] • Under case-control sampling, i.e., (ref), we have $\theta_\mathrm{AR}(x) = r_\mathrm{CC}(x,p_0) \Gamma_{\mathrm{CC},\mathrm{AR}}(x,p_0)$. • Under case-population sampling, i.e., (ref), we have $\theta_\mathrm{AR}(x) = r_\mathrm{CP}(x,p_0) \Gamma_{\mathrm{CP},\mathrm{AR}}(x,0)$. \end{enumerate}

Unlike the case of $\theta_\mathrm{RR}(x)$, $\theta_\mathrm{AR}(x)$ remains unidentified even in (ref). This happens because $\Gamma_{\mathrm{CP},RR}(x,0)$ is a ratio of two terms, where $r_\mathrm{CP}(x,p_0)$ cancels out, but $\Gamma_{\mathrm{CP},\mathrm{AR}}(x,0)$ is a difference and the common factor $r_\mathrm{CP}(x,p_0)$ does not disappear. Also, unlike $\theta_\mathrm{RR}(x)$, the rare disease approximation does not provide anything useful in either of the two sampling schemes: if $p_0\approx 0$, then $r_\mathrm{CC}(x,p_0) \approx 0$ and $r_\mathrm{CP}(x,p_0) \approx 0$ by continuity. However, the partial identification approach still remains useful.

theoremSuppose that (ref) are satisfied. For all $x\in \mathcal{X}$, we have the following. \begin{enumerate}[(1)] • Under case-control sampling, i.e., (ref), \[ \min_{p\in [0,\bar p]} r_\mathrm{CC}(x,p) \Gamma_{\mathrm{CC},\mathrm{AR}}(x,p)\leq \theta_\mathrm{AR}(x)\leq \max_{p\in[0, \bar p]} r_\mathrm{CC}(x,p)\Gamma_{\mathrm{CC},\mathrm{AR}}(x,p), \] where the bounds are sharp. • Under case-population sampling, i.e., (ref), \[ \min\bigl\{ 0,\ r_\mathrm{CP}(x,\bar p)\Gamma_{\mathrm{CP},\mathrm{AR}}(x,0)\bigr\} \leq \theta_\mathrm{AR}(x) \leq \max\bigl\{ 0,\ r_\mathrm{CP}(x,\bar p)\Gamma_{\mathrm{CP},\mathrm{AR}}(x,0)\bigr\}, \] where the bounds are sharp. \end{enumerate}

Since $\theta_\mathrm{AR}(x)$ is a difference of probabilities, it is always between $-1$ and $1$. Indeed, we show in the proof that all the bounds in (ref) lie within the interval between $-1$ and $1$. (ref) is a simple corollary of (ref): sharpness follows from the fact that $p_0$ is unidentified and that $r_\mathrm{CC}(x,p)\Gamma_{\mathrm{CC},\mathrm{AR}}(x,p)$ and $r_\mathrm{CP}(x,p) \Gamma_{\mathrm{CP},\mathrm{AR}}(x,p)$ are all continuous in $p$. Unlike the case of random sampling, the conditional average treatment effect is only partially identified even under strong ignorability. Also, it is noteworthy that in (ref), the sign of $\theta_\mathrm{AR}(x)$ is determined by that of $\Gamma_{\mathrm{CP},\mathrm{AR}}(x,0)$: if we know that $\theta_\mathrm{AR}(x)\geq0$, then we know that the conditional average treatment effect is at most $r_\mathrm{CP}(x,\bar p) \Gamma_{\mathrm{CP},\mathrm{AR}}(x,0)$.

We now consider replacing unconfoundedness with the monotonicity assumptions.

theoremSuppose that (ref) are satisfied. Then, for all $x\in \mathcal{X}$, we have the following. \begin{enumerate}[(1)] • Under case-control sampling, i.e., (ref), \[ 0 \leq \theta_\mathrm{AR}(x) \leq \max_{p\in[0,\bar p]} r_\mathrm{CC}(x,p) \Gamma_{\mathrm{CC},\mathrm{AR}}(x,p), \] where the bounds are sharp. • Under case-population sampling, i.e., (ref), \[ 0 \leq \theta_\mathrm{AR}(x) \leq r_\mathrm{CP}(x, \bar p) \Gamma_{\mathrm{CP},\mathrm{AR}}(x,0), \] the bounds are sharp. \end{enumerate}

Similarly to our comments below (ref), knowledge of $p_0$ is potentially useful to improve the bounds given in (ref): this point will be relevant when we discuss aggregation in the following section. This is so because, by the Bayes rule, the difference between the two prospective regression functions that appear in (ref) can be shown to be equal to $r_\mathrm{CC}(x,p_0)\Gamma_{\mathrm{CC},\mathrm{AR}}(x,p_0)$ under (ref) and to $r_\mathrm{CP}(x,p_0)\Gamma_{\mathrm{CP},\mathrm{AR}}(x,0)$ under (ref), respectively. However, $p_0$ is unrestricted in general, and hence maximizing over $p_0\in [0, \bar p]$ under (ref) delivers the sharp upper bounds.

The bounds in (ref) are comparable with those in (ref). In case-control or case-population sampling, strong ignorability is not as powerful as in random sampling. First, strong ignorability does not deliver point identification of the conditional average treatment effect. Second, the monotonicity assumptions do restrict the sign of $\theta_\mathrm{AR}(x)$, but, otherwise, they have the same amount of information as the strong ignorability assumptions in terms of the maximum admissible value of $\theta_\mathrm{AR}(x)$.

Aggregation

Conditioning on a specific value of the covariate vector and aiming at $\theta_\mathrm{RR}(x)$ or $\theta_\mathrm{AR}(x)$ as in (ref) is one natural approach to deal with potential heterogeneity in the causal treatment effect. However, the corresponding bounds as functions of $x$ (e.g., $\mathrm{OR}(x)$) are complicated objects, and they are difficult to estimate with high precision when $X^*$ is multi-dimensional.

To avoid the curse of dimensionality, it is popular in case-control studies to adopt logistic regression. Some authors have alternatively parametrized the odds ratio function itself in case-control studies, focusing on establishing a doubly robust estimator of the odds ratio: see e.g.\ yun2007semiparametric and tchetgen2013closed. Direct parametrization of $\Gamma_{d,\mathrm{AR}}(x,p)$ appears to be uncommon though.

Parametric assumptions are convenient, but they are restrictive: e.g.\ $\mathrm{OR}(x)$ is generally an unknown function of $x$ that can be highly nonlinear. Instead of introducing any parametrization, aggregation over the population distribution of the covariates can be a useful approach to obtain a robust summary measure.

If one wants to report an aggregated parameter such as $\int_\mathcal{X} \theta_\mathrm{AR}(x) \omega(x) dx$ for some weight function $\omega$, sharp bounds can be obtained by taking max/min over $p_0$ after aggregation. The most natual choice of the weight function $\omega$ is probably the true population density of $X^*$. The distribution of $X^*$ is unidentified in case-control studies, but the situation is not too bad because the only unidentified object is, again, $p_0$.

Consider the following aggregated parameters:

equation[equation omitted — 273 chars of source]

$\bar \vartheta_\mathrm{AR}$ is the standard average treatment effect. For $\bar \vartheta_\mathrm{RR}$, we use the logarithm of $\theta_\mathrm{RR}(X^*)$ to take an average. Since $\mathbb{E}\{ \log\mathrm{OR}(X^*)\} \leq \log\mathbb{E}\{ \mathrm{OR}(X^*)\}$ by Jensen's inequality, the average of the logarithm is less likely to be affected unduly by outliers. We also note that it is more conventional to work with the logarithm of the odds ratio than the odds ratio itself. If one still prefers aggregating $\theta_\mathrm{RR}(x)$ itself, it is straightforward to modify our methodology by using the same principle outlined in this section.

Our approach is to use the fact that the only missing piece in case-control or case-population samples is $p_0$. We first derive sharp identifiable bounds on $\theta_\mathrm{RR}(x)$ and $\theta_\mathrm{AR}(x)$ with $p_0$ given. We then aggregate over the distribution of $X^*$, which depends on $p_0$ in case-control studies. Specifically, we use the fact that for all $x\in \mathcal{X}$, \[ f_{X^*}(x) = \left\{

aligned&f_{X|Y}(x\mid 1) p_0 + f_{X|Y}(x\mid 0) (1-p_0) \quad &&in case-control studies,\\ &f_{X|Y}(x\mid 0) \quad &&in case-population studies.

\right. \] We can then rely on (ref) to address the fact that $p_0$ is unidentified. For this purpose, we can maximize or minimize over $p_0\in [0,\bar p]$ to obtain bounds, or, more informatively, we can plot the whole bound functions on $[0, \bar p]$: choosing the maximal value that is allowed for $\bar p$ (e.g., $\bar p = 1$ in case-control studies) corresponds to the case where we have no information for $p_0$. This line of reasoning leads to the main results in this section.

We will use the following objects: for $d\in \{\mathrm{CC},\mathrm{CP}\}$, \[ \Psi_{d,\mathrm{RR}}(p,y) := \mathbb{E}\{ \log\Gamma_{d,\mathrm{RR}}(X,p) \mid Y=y \}. \] The logarithm in the definition of $\Psi_{d,\mathrm{RR}}(p,y)$ is because $\bar\vartheta_\mathrm{RR}$ is the aggregation of $\log\theta_\mathrm{RR}(x)$. If one wants to bound $\int_\mathcal{X} \theta_\mathrm{RR}(x) f_{X^*}(x) dx$, then changing the definition of $\Psi_{d,\mathrm{RR}}(p,y)$ to $\mathbb{E}\{ \Gamma_{d,\mathrm{RR}}(X,p) \mid Y=y \}$ will do. Also, we note that $\int_\mathcal{X} \theta_\mathrm{RR}(x) f_{X^*}(x) dx$ differs from the ratio of unconditional counterfactual probabilities. Let \[

aligned\Psi_{\mathrm{CC},\mathrm{AR}}(p,y) &:= \mathbb{E}\{ r_\mathrm{CC}(X,p)\Gamma_{\mathrm{CC},\mathrm{AR}}(X,p) \mid Y = y \},\\ \Psi_{\mathrm{CP},\mathrm{AR}}(p) &:= \mathbb{E}\{ r_\mathrm{CP}(X,p) \Gamma_{\mathrm{CP},\mathrm{AR}}(X,0) \mid Y=0 \bigr\},

\] where we note that $\Psi_{\mathrm{CP},\mathrm{AR}}(p)$ is a simple linear function of $p$ by definition. Finally, for $k\in \{\mathrm{RR},\mathrm{AR}\}$, define $\mathcal{C}_{\mathrm{CC},k}(p)$ by a convex combination of $\Psi_{\mathrm{CC},k}(p,1)$ and $\Psi_{\mathrm{CC},k}(p,0)$: i.e., $\mathcal{C}_{\mathrm{CC},k}(p) := \Psi_{\mathrm{CC},k}(p,1)p + \Psi_{\mathrm{CC},k}(p,0)(1-p)$.

theoremSuppose that (ref) are satisfied. We then have the following. \begin{enumerate}[(1)] • Under case-control sampling, i.e., (ref), the sharp identified bounds on $\bar\vartheta_\mathrm{RR}$ and $\bar\vartheta_\mathrm{AR}$ are given by \begin{align*} \min_{p\in[0,\bar p]} \mathcal{C}_{\mathrm{CC},\mathrm{RR}}(p) &\leq \bar \vartheta_\mathrm{RR} \leq \max_{p\in[0,\bar p]} \mathcal{C}_{\mathrm{CC},\mathrm{RR}}(p),\\ \min_{p\in[0,\bar p]} \mathcal{C}_{\mathrm{CC},\mathrm{AR}}(p) &\leq \bar \vartheta_\mathrm{AR} \leq \max_{p\in[0,\bar p]} \mathcal{C}_{\mathrm{CC},\mathrm{AR}}(p). \end{align*} • Under case-population sampling, i.e., (ref), we have $\bar \vartheta_\mathrm{RR} = \Psi_{\mathrm{CP},\mathrm{RR}}(0,0)$, where we remark that this point identification result does not require (ref). Further, the sharp identified bounds on $\bar\vartheta_\mathrm{AR}$ are given by \[ \min\left\{ 0,\ \Psi_{\mathrm{CP},\mathrm{AR}}(\bar p) \right\} \leq \bar\vartheta_\mathrm{AR} \leq \max\left\{ 0,\ \Psi_{\mathrm{CP},\mathrm{AR}}(\bar p) \right\}. \] \end{enumerate}
theoremSuppose that (ref) are satisfied. Then, we have the following. \begin{enumerate}[(1)] • Under the case-control sampling, i.e., (ref), the sharp identified bounds on $\bar\vartheta_\mathrm{RR}$ and $\bar\vartheta_\mathrm{AR}$ are given by \begin{equation*} 0\leq \bar\vartheta_\mathrm{RR} \leq \max_{p\in [0,\bar p]} \mathcal{C}_{\mathrm{CC},\mathrm{RR}}(p) \quad and \quad 0\leq \bar\vartheta_\mathrm{AR} \leq \max_{p\in [0,\bar p]} \mathcal{C}_{\mathrm{CC},\mathrm{AR}}(p). \end{equation*} • Under the case-population sampling, i.e., (ref), the sharp identified bounds on $\bar\vartheta_\mathrm{RR}$ and $\bar\vartheta_\mathrm{AR}$ are given by \begin{equation*} 0 \leq \bar\vartheta_\mathrm{RR} \leq \Psi_{\mathrm{CP},\mathrm{RR}}(0,0) \quad and \quad 0 \leq \bar\vartheta_\mathrm{AR} \leq \Psi_{\mathrm{CP},\mathrm{AR}}(\bar p), \end{equation*} where we remark that the bounds on $\bar\vartheta_\mathrm{RR}$ do not rely on (ref). \end{enumerate}

Generally, in both cases of strong ignorability and monotonicity, case-population sampling provides an easier environment for causal inference than case-control studies: $\Psi_{\mathrm{CP},\mathrm{RR}}(0,0)$ does not depend on $p$ and $\Psi_{\mathrm{CP},\mathrm{AR}}(p)$ is linear in $p$. Also, the bounds under strong ignorability are all comparable with those under monotonicity: the upper bounds have the same form under strong ignorability as under monotonicity except that the monotonicity assumptions impose restrictions on the direction of the causal effect.

(ref) show that $\bar \vartheta_\mathrm{RR}$ suites better case-control or case-population studies than $\bar \vartheta_\mathrm{AR}$, especially when the case is potentially rare, despite the popularity of the latter in random sampling. Specifically, $r_\mathrm{CC}(x,p)$ and $r_\mathrm{CP}(x,p)$ should be taken into account for $\bar\vartheta_\mathrm{AR}$, but they are irrelevant for $\bar\vartheta_\mathrm{RR}$. This is an important difference because $r_\mathrm{CC}(X,0) = r_\mathrm{CP}(X,0) = 0$, which implies that the bounds on $\bar\vartheta_\mathrm{AR}$ cannot be tighter under strong ignorability than under monotonicity. In order to see the point more clearly, consider the case of case-control studies, i.e., (ref), and suppose that $\max_{p\in[0,\bar p]} \bigl\{ \Psi_{\mathrm{CC},\mathrm{AR}}(p,1)p + \Psi_{\mathrm{CC},\mathrm{AR}}(p,0)(1-p) \bigr\} > 0$ so that the upper bound on $\bar\vartheta_\mathrm{AR}$ is positive both under strong ignorability and under monotonicity. In this case, the lower bound on $\bar\vartheta_\mathrm{AR}$ under strong ignorability can never be strictly positive because $\Psi_{\mathrm{CC},\mathrm{AR}}(0,y)$ is trivially equal to zero. In other words, strong ignorability does provide a more informative environment than monotonicity but only in the sense that the former does not restrict the sign of $\bar\vartheta_\mathrm{AR}$. Once the sign of $\bar\vartheta_\mathrm{AR}$ is given, then there is nothing extra the strong ignorability assumptions offer relative to the monotonicity setup in understanding the average treatment effect. The same is true for the case-population case, i.e., (ref).

If we focus on $\bar\vartheta_\mathrm{RR}$, then the average of the log odds ratios, i.e., $\beta(y):= \Psi_{\mathrm{CC},\mathrm{RR}}(0,y) = \Psi_{\mathrm{CP},\mathrm{RR}}(0,y) = \mathbb{E}\{ \log \mathrm{OR}(X) \mid Y=y\}$ becomes the central object for estimation and inference. For instance, in (ref), all we need is $\beta(0)$, which can be interpreted as $\bar\vartheta_\mathrm{RR}$ itself or its sharp upper bound, depending on whether we assume strong ignorability or monotonicity, respectively. In (ref), if (ref) is imposed, then $\Psi_{\mathrm{CC},\mathrm{RR}}(p,y)$ can be shown to be decreasing in $p$, and therefore we have $\mathcal{C}_{\mathrm{CC},\mathrm{RR}}(p) \leq \beta(1)p + \beta(0) (1-p)$. Since the right-hand side is linear in $p$, we can easily conduct inference on $\bar\vartheta_\mathrm{RR}$ uniformly in $p\in [0,\bar p]$ by using $\beta(y)$, though this can be conservative.

The log odds ratio $\log\mathrm{OR}(x)$ has been a popular measure of association in case-control studies, and $\beta(y)$ is an aggregation of it by using the identified distribution of $X$ given $Y=y$. AAA establish the semiparametric efficiency bound for $\mathbb{E}\{ \log \mathrm{OR}(X) \}$ and suggest efficient estimators that accommodate high-dimensional machine learning estimators in the first stage. For low-dimensional $X$, a straightforward algorithm for efficient estimation of $\beta(y)$ is available, and it can be easily implemented using standard software. The algorithm is described in online appendix B2.

Causal inference under monotonicity

In this section, we discuss how to carry out causal inference on the aggregated parameters $\bar\vartheta_\mathrm{RR}$ and $\bar\vartheta_\mathrm{AR}$ under the MTR and MTS assumptions: inference under strong ignorability can be done by the same principles. In our discussion below, $z(1-\alpha)$ will be the $1-\alpha$ quantile of the standard normal distribution.

We first consider relative risk, for which we use $\exp(\bar\vartheta_\mathrm{RR})$ as the parameter of interest: see our discussion right below (ref). Our basis for inference is (ref). Let $\beta(y) := \Psi_{\mathrm{CP},\mathrm{RR}}(0,y) = \mathbb{E}\{ \log \mathrm{OR}(X) \mid Y=y\}$ for $y=0,1$.

Inference is easier when we have a case-population sample: all we need is $\beta(0)$. Since we have $1\leq \exp(\bar\vartheta_\mathrm{RR}) \leq \exp\{ \beta(0) \}$ by (ref), a $1-\alpha$ confidence interval for $\exp(\bar\vartheta_\mathrm{RR})$ can be constructecd by $\bigl[ 1,\ \exp\{ \hat\beta(0) + z(1-\alpha) \hat s(0) \} \bigr]$, where $\hat\beta(0)$ is an asymptotically normal estimator of $\beta(0)$, and $\hat s(0)$ is its standard error.

In the case of case-control sampling, i.e., (ref), we should base our inference on $\mathcal{C}_{\mathrm{CC},\mathrm{RR}}(\cdot)$. However, $\mathcal{C}_{\mathrm{CC},\mathrm{RR}}(p)$ is nonlinear in $p$, and hence it is difficult to obtain a confidence band uniformly in $p\in[0,\bar p]$. We propose two solutions. One is just to use one-sided pointwise confidence bands using bootstrap, akin to algorithm 2 in online appendix D, where we focus on pointwise inference for $\bar\vartheta_\mathrm{AR}$. The other is to take a conservative approach by using the fact that $\mathcal{C}_{\mathrm{CC},\mathrm{RR}}(p) \leq \tilde\beta(p) := \beta(1) p + \beta(0)(1-p)$. Specifically, in online appendix D, we show that

equation[equation omitted — 195 chars of source]

where $u(1-\alpha) := z(1-\alpha/2) \max\{ \hat s(0),\ \hat s(1) \}$ with $\hat s(y)$ is the standard error of $\hat \beta(y)$, the asymptotically normal estimator of $\beta(y)$.

We now turn to inference on $\bar\vartheta_\mathrm{AR}$. The case-population sample provides an easier environment again: we can exploit the fact that $\Psi_{\mathrm{CP},\mathrm{AR}}(p)$ is a simple linear function with the form of $\Psi_{\mathrm{CP},\mathrm{AR}}(p) := p \xi_\mathrm{CP}$, where $\xi_\mathrm{CP}$ is implicitly defined here and does not depend on $p$. For more details, see online appendix D.

Inference on $\bar\vartheta_\mathrm{AR}$ with a case-control sample relies on the function $\mathcal{C}_{\mathrm{CC},\mathrm{AR}}(\cdot)$, and its nonlinearity in $p$ makes it difficult to construct a uniform confidence band. Since $\Psi_{\mathrm{CC},\mathrm{AR}}(\cdot,y)$ is not monotonic, the conservative approach we discussed for $\bar\vartheta_\mathrm{RR}$ does not apply here. Therefore, we propose using one-side pointwise confidence intervals, for which we use Efron's bias-corrected percentile intervals. Computational details for implementation are given in online appendix D.

Empirical Examples

Case-control sampling: entering a very selective university

We consider quantifying the causal effect of attending private school on entering a very selective university by using the Pakistan data collected by Delavande:Zafar:19. This is survey data from male students who were already enrolled in different types of universities in Pakistan, all located in Islamabad/Rawalpindi and Lahore. Delavande:Zafar:19 include two Western-style universities, one Islamic university, and four madrassas, but we focus on the two Western-style ones in our analysis: between the two universities, Delavande:Zafar:19 call the more expensive, selective, and reputable university "Very Selective University" (VSU) and the other simply "Selective University" (SU). Therefore, we restrict the population of interest to those who entered either VSU or SU, and we define the binary outcome to be whether a student entered VSU. The binary treatment we consider is whether a student attended private school before university. Since the students in the sample were already enrolled in either VSU or SU at the time of the survey, we have a case-control sample, i.e. our (ref).

table[table omitted — 498 chars of source]

(ref) shows the likelihood of entering VSU by private school attendance before university. The empirical odds ratio is 1.38.

In this example, the unconfoundedness assumption is unlikely to hold, because those who attended private school before university are likely to have more resourceful parents. This concern may not completely disappear even if we control for parental income and wealth because of the presence of unobserved parental abilities and resources that could affect their children's university choice. However, the MTR and MTS assumptions are still plausible: private school is probably no inferior input to university preparations (hence, MTR), and those who actually chose to attend private school probably care about their future college choice no less than those who did not (hence, MTS). Then, the odds ratio of 1.38 can be interpreted as a sharp upper bound on causal relative risk; therefore, the effect of attending private school seems, at best, modest.

Now, we consider controlling for family background variables. Specifically, we include an indicator for at least one college-educated parent and parents' monthly income as covariates. (ref) reports estimation results for the aggregated log odds ratio within each of the case and the control: both $\beta(y)$ and $\exp\{\beta(y)\}$ convey the same information, but $\exp\{ \beta(y)\}$ is easier to interpret because it is comparable to the usual odds ratio in terms of its scale. The fact that $\widehat\beta(1)$ and $\widehat\beta(0)$ are notably different suggests that the amount of heterogeneity among individuals may be substantial. The confidence intervals are computed based on the MTR and MTS: hence, they are one-sided.

table[table omitted — 588 chars of source]
figure[figure omitted — 641 chars of source]

We now consider the methods described in (ref), i.e., causal inference on the aggregated relative and attributable risk (RR and AR, respectively) in terms of the population distribution of the covariates. We rely on the MTR and MTS assumptions to interpret our results as upper bounds. For AR, we use the same covariates as in RR. The number of the bootstrap replications was 10,000.

(ref) summarizes the results: the left (right) panel shows RR (AR). The case probability $p_0$ of entering VSU in the population is not identified in this dataset. But we can trace out the upper bounds as the value of $p_0$ varies between $0$ and $1$.

Consider the left panel of (ref), i.e., RR, where we take the conservative approach and plot $\tilde\beta(p) = \beta(1) p + \beta(0)(1-p)$. If we take the point estimate at face value, attending private school increases the chance of entering VSU by a factor of at most 1.26. Even in terms of the confidence intervals, it seems highly unlikely that the impact is more than a factor of 2. The right panel of (ref) shows AR. The graph shows an inverted U-shape, because $r_\mathrm{CC}(x,p)\Gamma_{\mathrm{CC},\mathrm{AR}}(x,p) = 0$ whenever $p$ is either $0$ or $1$. The maximum point estimate of the upper bound is 0.044, while the maximum value of the confidence intervals is 0.153. Therefore, it seems highly unlikely that attending private school increases the chance of entering VSU by more than 16 percent.

None of our results require strong ignorability or the rare-disease assumption. Our conclusion of a relatively small positive effect of attending private school on entering VSU, if it exists at all, is reminiscent of existing results in labor economics that find access to private schools have only modest effects on children's performance Epple:17,MacLeod:19.

Case-population sampling: joining a criminal gang

We revisit carvalho2016living, who combine the 2000 Brazilian Census with a unique survey of drug-trafficking gangs in favelas (slums) of Rio de Janeiro; therefore, their dataset is an example of case-population sampling, i.e., our (ref). In their study, they use the method of lancaster1996case to estimate a model of selection into the gang by using race, age, illiteracy, house ownership, and religiosity. They note that the five characteristics are likely to be predetermined while years of schooling may be endogenous to entry, i.e., joining the gang may lead members to drop out of school. Indeed, 90 percent of gang members are not in school, whereas 46 percent of men aged 10--25 are not in school. They find that “younger individuals, from lower socioeconomic background (black, illiterate, and from poorer families) and with no religious affiliation are more likely to join drug-trafficking gangs.”

(ref) provides summary statistics of the sample. We regard currently not attending school as the treatment variable of interest. Unconfoundedness is not plausible because of the endogeneity of schooling that we mentioned earlier. Furthermore, it is plausible that unmeasured factors such as family support could affect both treatment and outcome. However, not being in school may increase the chance of exposure to gang-related activities, and those who chose to be in school may be the ones who care for consequences no less than those who chose not to be; therefore, the MTR and MTS follow, respectively.

table[table omitted — 747 chars of source]

(ref) presents estimation results for $\beta(y)$ and $\exp\{\beta(y)\}$, for which we control for the same covariates as carvalho2016living. Unlike (ref), $y=0$ now corresponds to the entire population. Therefore, $\beta(0)$ itself is the log odds ratio aggregated over the population, which is the sharp upper bound on the aggregation of the log causal relative risk, i.e., $\bar\vartheta_\mathrm{RR}$. The point estimate of $\beta(0)$ is 2.71, and that of $\exp\{\beta(0)\}$ is 15.01, which suggests that the chance of those who are not in school joining a gang may be (up to) 15 times as large as that of those who are.

table[table omitted — 662 chars of source]

Our discussion above can be supplemented by checking the causal AR. At the three points of $p_0\in\{0.05,0.10, 0.15\}$ that carvalho2016living considered, the point estimates and the end-points of the uniform confidence interval (in parentheses) for the upper bound on the causal AR are 0.33 (0.43), 0.66 (0.86), and 0.99 (1), respectively. The uniform confidence band is based on 1,000 bootstrap replications. Note that the confidence band is truncated at one, because AR cannot be larger than one.

Overall, our results are suggestive of potentially large impacts of keeping young men in school in order to discourage them to participate in criminal activities. Further research based on careful study designs would be necessary to reach a more definitive answer.

Random sampling: physician's hours

FangGong17 construct estimates for physicians’ hours spent on Medicare beneficiaries and find that about 3 percent of physicians billed for more than 100 hours per week. They refer to these physicians as flagged physicians. FangGong17 state that “flagged physicians are slightly more likely to be male, non-MD, more experienced, and provide fewer E/M services. Importantly, they work in substantially smaller group practices (if at all), and have fewer hospital affiliations.” We use their study to illustrate the findings in this paper. Specifically, the outcome variable is whether a physician billed for more than 100 hours per week in either 2012 or 2013, the treatment variable is a binary indicator showing whether or not the number of group practice numbers is less than 6, which is the median size in the data, and the covariates include an indicator for male, an indicator for doctor of medicine (MD), and experience in years (cubic polynomial).

table[table omitted — 560 chars of source]

The original dataset in FangGong17 is updated in FangGong20 after Matsumoto pointed out data and coding errors in the original work. In our analysis, we use the updated dataset. (ref) summarizes the sample, which consists of 78,165 physicians who billed at least 20 hours per week.

Treating this dataset as a random sample, we extract a case-control dataset: the case sample is composed of 2,261 flagged physicians; the control sample of equal size is randomly drawn without replacement from the pool of physicians who were never flagged. Analogously, a case-population dataset is obtained by combining the case sample with a population sample of equal size that is randomly drawn without replacement from all observations and its flagged status is coded missing.

It is highly unlikely that the group practice size is exogenous conditional on a small number of covariates; hence, we rely on the monotonicity assumptions (i.e., MTR and MTS) and focus on the upper bounds on the relative and attributable risks.

table[table omitted — 769 chars of source]

(ref) reports the estimates of $\exp[\beta(y)]$ and their one-sided confidence intervals of $\exp[\beta(y)]$ for each sampling scheme. The lower bound is 1 because of the MTR assumption and averaging is done for a relevant population in each column. Because the proportion of the flagged physicians is less than 0.03, we invoke the rare disease assumption and regard $\exp[\beta(y)]$ as an approximation to the upper bound on RR. Although there are some noticeable differences across different columns, the estimates are similar. This is consistent with the identification result that the price to pay is less for identification of RR when we move from random sampling to outcome-dependent sampling. The story is different if we focus on AR. (ref) shows that the upper bounds on AR under outcome-dependent sampling are much larger than those under random sampling, especially when $\bar p = 0.1$.

table[table omitted — 777 chars of source]

\setstretch{1} \setstretch{1.5}