EconBase
← Back to paper

Formalising causal inference as prediction on a target population

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,526 characters · 18 sections · 74 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Formalising causal inference as prediction on a target population

abstractThe standard approach to causal modelling especially in social and health sciences is the potential outcomes framework due to Neyman and Rubin. In this framework, observations are thought to be drawn from a distribution over variables of interest, and the goal is to identify parameters of this distribution. Even though the stated goal is often to inform decision making on some target population, there is no straightforward way to include these target populations in the framework. Instead of modelling the relationship between the observed sample and the target population, the inductive assumptions in this framework take the form of abstract sampling and independence assumptions. In this paper, we develop a version of this framework that construes causal inference as treatment-wise predictions for finite populations where all assumptions are testable in retrospect; this means that one can not only test predictions themselves (without any fundamental problem) but also investigate sources of error when they fail. Due to close connections to the original framework, established methods can still be be analysed under the new framework.

Introduction

For many problems in the social and health sciences, it is important to analyse and predict the efficacy of treatments or policies. This requires causal modelling, as observational outcome distributions may not reflect outcome distributions under active treatment policies. The Potential Outcome framework due to Neyman neyman1923 and rubin1974 is the dominant approach to causal modelling in many fields. Despite Neyman's early work, we shall follow holland1986 in calling these models Rubin causal models (RCMs), as we often specifically refer to Rubin's now-dominant version that is framed in terms of probability distributions. RCMs are the preferred framework in particular for informing specific policies or interventions, due to the focus on specific outcomes of interest and the aptitude to accommodate individual problem settings imbens2020, markus2021 -- vis-a-vis the standard econometric approach heckman2024 and structural equation models pearl2009, which are are sometimes seen as having the edge in more abstract theory building, including the modelling of unobservables. However, while the `credibility revolution’ of the last decades resulted in theoretically rigorous methods angrist2010, it `has focused primarily on internal validity’ egami2023. In contrast, while already the original paper introducing RCMs called on `investigators [to] carefully describe their sample of trials and the ways in which they may differ from those in the target population' rubin1974, the analysis of external validity has been widely neglected (as attested by lamentations across fields, see Section (ref)).

We argue that this has in part to do with the framework itself, as it does not provide a straightforward way to model target populations. Instead of effects on target populations, the focus of the framework is the `identification' of parameters in some abstract probability distribution -- `as if a parameter, once well established, can be expected to be invariant across settings’ deaton2018. A further issue with this abstract framing is that it is very difficult, if possible at all, to formulate concrete and testable assumptions that enable the accurate prediction of the effects of policies on target populations. In this paper, we suggest an amendment to the framework to overcome, or at least mitigate, these problems. More concretely, we suggest a variant of RCMs that directly models both observed and target populations and their relationship, avoiding detours through abstract distributions. It shifts the focus from the identification of true parameters to the prediction of future outcomes based on (retrospectively) testable assumptions. Rather than ignoring existing causal inference methodology, we show that the new framework can capture established estimators and provide a complementary perspective on them. We, thus, provide an `intermediary' framework that establishes links between high-level intuitions, as formalised in RCMs, on the one hand and directly testable assumptions about concrete populations on the other. Hence, this variant augments the strengths of RCMs for evidence-based policy making by focussing on predictions for concrete target populations grounded in testable assumptions. Beyond these benefits, it also offers complementary perspectives not only on established causal estimators but also on causal inference as a whole.

The structure of this work is as follows. In Section (ref), we recapitulate RCMs and discuss calls across fields to direct more attention to problems with external validity and to model the target population more directly. In Section (ref), we introduce the new framework in the context of simple estimators; a broader survey of existing estimators and how they fit into the new framework can be found in Appendix (ref). In Section (ref), we compare the new with the conventional framework, by drawing formal connections and discussing differences in practice. While the bulk of the paper focuses on average treatment effects and potential outcomes, Section (ref), briefly addresses generalisations to conditional treatment rules as well as to distributional properties beyond the mean. Section (ref) concludes and discusses how the framework how the framework accommodates a less metaphysically loaded view on causal inference as a whole.

Background: Rubin Causal Models (RCMs)

In this section, we first introduce the standard formalism of RCMs, before critically discussing the assumption of an underlying probability distribution as well as the problem of explicitly modelling outcomes on a target population.

The framework

The main components of RCMs are the following random variables (RVs): $T_i$ is the decision variable indicating whether person $i$ is treated and takes values in $\{0,1\}$, which represent control and treatment.\footnote{We only consider the binary treatment case here as this makes the presentation of RCMs simpler.} $Y_{1i}$ and $Y_{0i}$ denote the outcome for $i$ upon receiving treatment and control, respectively. Based on this, we can define the actual outcome

equation[equation omitted — 100 chars of source]

In the example of job trainings, $T_i$ indicates whether someone gets offered job training and $Y_i$ indicates whether they have a job after a fixed time, say, one year. It is, thus, assumed that for every person, both $Y_{0i}$ and $Y_{1i}$ are well-defined, i.e., whether they find a job if they don't get offered job training and whether they find a job if they are assigned job training, respectively. Of course, we can, in principle, only observe the value of one of the two variables for each individual.

In many settings (some of which we will consider in this paper), we also have informative covariates $X_i$. In the context of job trainings, these may be attributes such as age, gender, education, and employment history. A foundational assumption behind RCMs is that there is a joint distribution $P$ over all variables, i.e.

equation[equation omitted — 74 chars of source]

We are then usually interested in the average treatment effect (ATE) $\mathbb{E}_P[Y_{1i} - Y_{0i}]$. The ATE is typically seen as the expectation over the individual treatment effect (ITE) $Y_{1i} - Y_{0i}$. The fact that we can only ever measure one of them for each $i$ has been dubbed the `fundamental problem of causal inference'.

In line with many textbooks, we showcase RCMs in the context of Randomised Controlled Trials (RCTs). This represents the `experimental ideal' in the sense that it assumes we have access to data from a randomised experiment. We can then assume that the potential outcomes are independent of the treatment decision

equation[equation omitted — 77 chars of source]

This means we don't need covariates $X_i$ here, as the ATE can be expressed as

equation[equation omitted — 120 chars of source]

and both terms on the RHS can directly be estimated from our data: Assuming that past and future data are sampled from the distribution $P$, the law of large numbers says that the empirical estimate of ((ref)) converges to the ATE almost surely.

RCTs is are often not available and many methods have been devised to allow causal inference when random assigment and thus the conditional independence ((ref)) is violated. Most of these methods rely on another assumption called unconfoundedness (also conditional independence, ignorability, selection-on-observables), given by

equation[equation omitted — 79 chars of source]

This means that treatment assignment may have been based on $X_i$ but not on any other variables that could give information about the outcome. In other words, the treatment group should be comparable to the control group if we take the covariates into account. For example, in labour market programmes, information about potential participants is taken into account for deciding who gets offered job training. The unconfoundedness assumption is is then assumed to be satisfied if the decision was only based on the covariates $X_i$ -- hence the name `selection on observables'.

The distribution

As described above, RCMs are formulated in terms of joint distributions that represent the ground truth and are used to derive theoretical guarantees. But how should these distributions be understood? Should we take them at face value and believe that they are supposed to correctly describe some data-generating process? Or are they just models that have a more indirect relationship with what we believe to be true about the world? We now discuss these two options in turn.

The first option is that descriptions in terms of generative distributions can be true or false. Much of the literature suggests this reading. This would mean e.g. that there is a true (marginal) distribution of covariates from which people are sampled. And this seems to hold then for any selection of covariates/attributes that we can come up with. In principle, the number of possible descriptions (choices of covariates) with is almost unlimited, although we often just take the attributes that we can most easily measure. The invoked distributions are commonly not thought to pertain to a specific time, though sometimes to a specific location. But the relationships between e.g. unemployment, age, and education clearly change over time -- if they can be said to be stable at any given point in the first place. Every case of a person finding or not finding a job is highly individual and it seems rather strong that they just follow a general law plus some noise -- a notion we know from physics. Nancy cartwright1999 makes the point that while a lot of work in physics has focused on finding latent quantities that do have stable relationships (like forces in Nowton's and Hooke's laws), social sciences like economics typically consider more easily measurable quantities that are particularly salient.\footnote{This applies particularly to RCMs which, in contrast to other causal modelling approaches heckman2024 do not model unobservables.} And `to suppose that there really is some probability measure over [such quantities], you need a lot of good arguments’ (p. 325).

The other option is, then, to say that aspects like sampling from a joint distributions are mere models that are always wrong (if they have truth conditions at all). On this view, assumed joint distributions are idealisations and aim to capture patterns that we can observe between individual events, but they cannot be literally true. Arguments of this kind go back at least to de Finetti, who showed that Bayesians with certain priors can equivalently describe their subjective credences as if they were sampling i.i.d. from imaginary distributions. If this is to be the interpretation, it is surprising how little attention has been paid to how the idealised model relates to the world we observe. What, for example, does the assumption of unconfoundedness mean if it can never be true? If generative distributions are useful fictions, under which conditions are they useful? While a realist about distributions should already provide bridges to observable statements of interest (beyond the invocation of i.i.d. samples), an instrumentalist has all the more reason to do so.

While a standard and seemingly innocuous aspect, these distributions actually do a lot of the heavy lifting. They provide the connections between both the quantities of interest and between different populations, via sampling assumptions. Getting the former right requires internal validity while getting the latter right requires external validity; in RCMs, these two aims are typically considered in separation. Internal validity is concerned with the assumptions discussed above, especially with whether unconfoundedness holds in the assumed `generative process' behind observational data. This is supposed to ensure that estimations of quantities like the ATE are correct for the observed sample. External validity, in contrast, concerns the question whether these estimations also hold for future data, that is, for populations on which we want to make treatment decisions in the future. It is usually assumed that the distribution reflected in our historical data is the same as or similar to the distribution from which future data is `sampled', which allows us to predict outcomes or treatment effects in some population of interest. Figure (ref) provides a visual sketch of the RCM picture, where the top and left parts (distribution and observed population) cover internal validity.

figure[figure omitted — 165 chars of source]

The target population

‘For almost any study to be of interest, the results must be generalizable to a population of trials.’ This quote comes from Rubin's first paper introducing the framework rubin1974. As he also noted,

quote`in order to generalize the results of any experiment to future trials of interest, we minimally must believe that there is a similarity of effects across time and more often must believe that the trials in the study are “representative” of the population of trials. [...] Even though the trials in an experiment are often not very representative of the trials of interest, investigators do make and must be willing to make this assumption [...] in order to believe their results are useful.’ rubin1974.

However, it has been often noted that `[s]ocial scientists frequently invoke external validity as an ideal, but they rarely attempt to make rigorous, credible external validity inferences.’ findley2021. It is striking how pervasive such statements are across various fields in which RCMs are used. In economics, `[r]esearchers tend to focus primarily on threats to internal validity' bo2021 whereas external validity is virtually not discussed in standard textbooks like angrist2010. In epidemiology, `threats to external validity are less well-understood’ as it is considered `the common view that external validity is secondary to, or contingent on, internal validity' lesko2020. In public health, `[t]he consequence of this emphasis on internal validity has been a lack of attention to and information about external validity, which has contributed to our failure to translate research into public health practice' steckler2008. In behavioral medicine, it has been observed that `[t]he majority of intervention studies conducted and reported in Annals [of Behavioral Medicine] and other health journals [...] are usually silent on external validity’ glasgow2006. Similarly, `Only 11% of all experimental studies and 13% of all observational causal studies published in the American Political Science Review from 2015 to 2019 contain a formal analysis of external validity in the main text, and none discuss conditions under which generalization is credible’ egami2023. This lack of engagement with external validity across fields is problematic, especially given that in common settings, internal validity in a sense `carries much less information than the external validity assumptions’ breskin2019.

We believe that the RCM framework and its focus on identification strategies has contributed to this development; as noted by egami2023, `the credibility revolution [...] has focused primarily on internal validity’. One reason for this arguably lies in the very notion of `identification' of parameters that define the underlying distribution -- `as if a parameter, once well established, can be expected to be invariant across settings’ deaton2018.\footnote{Note however, that this problem is not unique to RCMs; indeed, advocacy of quick generalization `from the actual study experience to the abstract, with no referent in place or time’, miettinen1985 predates its widespread adoption e.g. in epidemiology, as criticized in keiding2016.}

Another reason is that in Rubin's models, it is difficult to follow Rubin's call for `investigators [to] carefully describe their sample of trials and the ways in which they may differ from those in the target population' rubin1974, given the difficulty to connect such a target population with the observed data in the formalism. The easiest setting would be to assume that the target population is sampled i.i.d. from the same distribution as the observed population. In this case, one can draw connections between them through the conjured distribution, using finite sample theory to estimate quantities of interest and possible errors or confidence intervals. This is, however, a very strong assumption; to say that the target population has a somewhat different generative process would mean to conjure a second abstract distribution that is somehow related to the first, from which the target population is then sampled. Making the relations between them explicit is difficult since none of the relations observed sample -- first distribution -- second distribution -- target population is observable and no assumptions about them can be tested.

Given the observed lack of engagement with external validity, some researchers have recently started to argue for more explicit modelling of the target population. For example, while Rubin acknowledges the vagueness of his notion of `representative' through quotation marks in the long quote above, rudolph2023 argue that `[a]ll statements regarding representativeness should make clear the way in which the study results generalise, the target population the results are being generalised to, and the assumptions that must hold for that generalisation to be scientifically or statistically justifiable’. Similar points have also been made by westreich2019 and fox2022. In the following, we propose an amendment to RCMs which allows to directly model the target population and allows to connect it to observed data without any detour through abstract distributions. This incentivises the engagement with concrete conditions that allow generalization and makes assumptions explicit and, retrospectively, testable. Still, the proximity and direct relation to RCMs makes it possible to keep trained intuitions and methodologies.

New framework

Setup and notation

We now introduce the setup and notation of the new version of the framework; it is close to the notation in manski2004. By $\mathcal{X}$, $\mathcal{Y}$, and $\mathcal{T}$ we denote the sets of possible covariates, outcomes (in $\mathbb{R}$), and treatments, respectively. We only consider binary treatments $\mathcal{T} = \{0,1\}$ in the main text for easier comparison with RCMs, although nothing changes for our framework with multiple treatments We assume that we have datapoints $(x_i, y_i, t_i)_{i \in \mathcal{J}}$ where $\mathcal{J}$ serves as the index set of our training data. That is, for each unit $i \in \mathcal{J}$ with covariates $x_i \in \mathcal{X}$, we have observed outcome $y_i \in \mathcal{Y}$ under treatment $t_i \in \mathcal{T}$.\footnote{This means that there is no notation for counterfactual outcomes of observed datapoints.}

We then consider a target population $\mathcal{I}$ with an unknown outcome function\footnote{manski2004 calls this the `response function'.}

equation[equation omitted — 79 chars of source]

such that $\mathsf{y}(i,t)$ denotes the outcome when treatment $t \in \mathcal{T}$ is assigned to individual $i \in \mathcal{I}$.\footnote{The common SUTVA assumption precluding interactions between treatment assignments to different individuals is encoded in the fact that the outcome only depends on the individual's treatment. One could allow such interactions by taking as inputs $i$ and a treatment vector of length $|\mathcal{I}|$.} Many methods for non-experimental settings require covariates; we denote covariates of target units by $\mathsf{x}(i), i \in \mathcal{I}$ using a representation function

equation[equation omitted — 60 chars of source]

Using a function for this highlights both that the covaruates of the target population may yet be unknown and that representing an individual $i$ through some covariates in a space $\mathcal{X}$ is itself a deliberate action.

As mentioned above, the most common quantity of interest is the average treatment effect (ATE); the finite-population version in our setting would be

equation[equation omitted — 176 chars of source]

This theoretical quantity is the difference between the average outcomes of assigning either treatment or control to everyone\footnote{We generalise this to more complex treatment rules in Section (ref).}, the average potential outcome (APO)

equation[equation omitted — 123 chars of source]

Our framework frames assumptions and results in terms of observable quantities that can be directly compared. To this end, we introduce a shorthand notation for approximate equality:

thm\begin{definition} For $\epsilon > 0$, we say that two values $r,s \in \mathbb{R}$ are $\epsilon$-similar if $|r - s| < \epsilon$. We write this as \begin{equation} r \approx_{\epsilon} s. \end{equation} \end{definition}

In the next section, we discuss fundamental assumptions that allow us to make predictions about the target population based on observed data.

Experimental setting: RCTs

In the comparatively simple case of RCTs, we do not need covariates for predicting the APO. Here, we individuals are randomly assigned into treatment and control group. In the earlier example, this could mean that a lottery is used to decide who is offered job training in a population of unemployed people. We denote treatment ($t=1$) and control ($t=0$) group in the observed data $\mathcal{J}$ by $\mathcal{J}_1$ and $\mathcal{J}_0$, that is,

equation[equation omitted — 118 chars of source]

The random assignment in RCTs can be used to justify the assumption that the average outcome in $\mathcal{J}_t$ approximates our quantity of interest ((ref)), that is, the average outcome in the target population if we assign treatment $t$ to everyone:

equation[equation omitted — 191 chars of source]

For this, we do not need to invoke any true distribution; it is enough to assume that the partition into observed and target samples as well as the partition into control and treatment groups can be considered random. That is, if the partition into $\mathcal{I}$ and $\mathcal{J}$ is random and the partition of $\mathcal{J}$ into $\mathcal{J}_0$ and $\mathcal{J}_1$ is random, then $\mathcal{J}_t$ under treatment $t$ is representative\footnote{rudolph2023 argue for such a notion of representativeness that is tied to a target population.} of $\mathcal{I}$ under treatment $t$: This means we treat them as being drawn from the same urn, similar to Neyman's original work.\footnote{It has been argued that such an urn model `applies rather neatly to the as-if randomized natural experiments of the social and health sciences' freedman2006.} This justifies ((ref)) because (for large enough data sets) the vast majority of possible partitions will lead to roughly equal averages in both parts. This can be shown with a Hoeffding inequality for finite samples (as in Proposition 1.2 of bardenet2015), giving statements of the form `for 95% of partitions, the difference between means is below 0.05'.

Such a measure on the number of admissible partitions can be made into a probabilistic statement (`with 95% probability...') if we additionally assume that all partitions are equally which is made in RCMs in the form of the i.i.d. assumption and in Bayesian frameworks such as dawid2021 in the form of exchangeability. In a finite population framework, this can, however, be neatly generalised to allow biased sampling schemes meng2018, meng2022; we elaborate on the connection to non-probability sampling in Appendix (ref). Our approach thus highlights this equi-probability assumption and allows to easily incorporate deviations, in the form of correlations between outcome and group membership. In our running example, this could take the form of incorporating e.g. fluctuations in general unemployment if the data collection happened a few years earlier.

Predictions and assumptions

Causal inference becomes more complicated outside controlled experiments and most literature is concerned with observational or quasi-experimental settings. For example, the data we have about the efficacy of job trainings typically does not come from RCTs. In such settings, we incorporate additional information in the form of covariates $x \in \mathcal{X}$. As we will demonstrate, these covariates are used not to analyse the difference between treatment and control groups, but between the treatment/control parts of the observed sample on the one hand and the full target sample on the other.

The last important concept in our framework is a predictor

equation[equation omitted — 70 chars of source]

As we will show now and, more extensively, in Appendix (ref), different estimators can be characterised by different predictors $p$, all with the goal of estimating the APO on the target sample though the average prediction on the observed sample. To make this work, two inductive assumptions are needed that tie the observed sample (and its treatment-wise groups) to the target sample (Figure (ref)).

figure[figure omitted — 471 chars of source]

This dissolves the traditional separation into internal and external validity (which should be seen as an advantage, as we will argue in the next section).

For external validity, the standard framework typically assumes that the target data look like the observed data in the sense that they are sampled from the same marginal distribution of $X_i$; we require a more specific and testable property: We assume that the average prediction on the observed population indexed by $\mathcal{J}$ is roughly the same as on the target population indexed by $\mathcal{I}$. We formalise this for $\epsilon > 0$ as the $\epsilon$-stable average predictions ($\epsilon$-SAP) assumption

equation[equation omitted — 95 chars of source]

where we denote the average prediction on observed and target sample by

equation*[equation* omitted — 103 chars of source]

and

equation*[equation* omitted — 114 chars of source]

respectively. While the assumption explicitly relates to the predictor $p$, it follows from the conventional assumption that covariates in sample and target population are sampled from the same distribution, at least for bounded $p$ and for large enough populations (see Section (ref)).

In addition to $\epsilon$-SAP, we need our predictor to work well on the target population. This assumption is formalised as $\delta$-calibration on the target population ($\delta$-CTP),

equation[equation omitted — 90 chars of source]

Essentially, this says that the errors of our predictor, when applied to the target population, will roughly average out. Together, these two assumptions allow us to make an $\epsilon+\delta$-good approximation of the APO under treatment $t$:

equation[equation omitted — 124 chars of source]

Different approaches to causal inference provide different estimators of the APO $\mu(\mathcal{I}, t)$ through different predictors $p$, for which the $\delta$-CTP assumption needs to be justified individually. Typically, $p$ is calibrated on $\mathcal{J}_t$ for all $t \in \mathcal{T}$ by design, so that $\delta$-CTP follows if we assume that the calibration error of $p$ on $\mathcal{I}$ is $\delta$-close to the calibration error on $\mathcal{J}_t$:

equation[equation omitted — 158 chars of source]

For example, we can see RCT-based inference as a using the degenerate predictor that predicts the group-wise mean for each individual, i.e.

equation[equation omitted — 110 chars of source]

Then

equation[equation omitted — 108 chars of source]

s.t. $\delta$-CTP then boils down to the assumption that the mean outcome in $\mathcal{J}_t$ is $\delta$-close to the mean outcome in $\mathcal{I}$ under treatment $t$, which we have discussed above. A more interesting case is matching or inverse probability weighting.

Observational setting and matching

One of the most basic techniques is exact matching rosenbaum1983. To capture this, define subgroups $\mathcal{I}^x, \mathcal{J}^x, \mathcal{J}_t^x$ for $x \in \mathcal{X}, t \in \mathcal{T}$,

align[align omitted — 193 chars of source]

For a sufficiently coarse-grained set of covariates, we may observe all combinations $(x,t)$ of covariates and treatments -- which is often called `positivity' or `common support':

equation[equation omitted — 117 chars of source]

Exact matching then uses the point-wise average predictor

equation[equation omitted — 119 chars of source]

This predictor is much more fine-grained that the RCT predictor ((ref)), as $p(x,t)$ predicts the observed $x$-wise average outcome in group $\mathcal{J}_t$, rather than averaging over all of $\mathcal{J}_t$. As we show in Proposition (ref), this predictor satisfies $\delta$-CTP ((ref)) if the average (signed) difference between the $x$-wise mean outcomes

equation*[equation* omitted — 232 chars of source]

is not strongly biased above or below zero, that is,

equation[equation omitted — 211 chars of source]

Note that in the limit of infinite data drawn from some distribution, the common unconfoundedness assumption ((ref)) would imply that the term $\mu_x(\mathcal{I}, t) - \mu_x(\mathcal{J}, t)$ goes to zero for every $x$ with probability one; this is strictly stronger than ((ref)), as the latter allows that the differences for different $x$ values cancel each other out (see Section (ref)). We now show that in practice, ((ref)) is indeed sufficient for the exact matching predictor; using the exact matching predictor for predicting the APO then amounts to inverse probability weighting:

thm\begin{proposition}[Predicting the APO through matching] \ \\ Fix any $t \in \mathcal{T}$. Assuming $\epsilon$-SAP ((ref)) for $p$ as in ((ref)) and average (signed) difference below $\delta$ as in ((ref)), the exact matching predictor ((ref)) gives us an $(\epsilon+\delta)$-good approximation of the APO $\mu(\mathcal{I}, t)$: \begin{equation} \left| \mu(\mathcal{I}, t) - \frac{1}{|\mathcal{J}|} \sum_{i \in \mathcal{J}_t} \frac{y_i}{e_t(x_i)} \right| < \epsilon + \delta, \end{equation} where $e_t(x) := \frac{|\mathcal{J}_t^x|}{|\mathcal{J}^x|}$ is the observed propensity score for treatment $t$. \end{proposition}
proofFirst, we get $\delta$-CTP for predictor from ((ref)) via \begin{align} \mu(\mathcal{I}, t) - \rho(\mathcal{I}, t) &= \frac{1}{|\mathcal{I}|} \sum_{i \in \mathcal{I}} \mathsf{y}(i, t) - p(\mathsf{x}(i),t) \\ &= \frac{1}{|\mathcal{I}|} \sum_{x \in \mathcal{X}} \sum_{i \in \mathcal{I}^x} \left(\mathsf{y}(i,t) - \sum_{j \in \mathcal{J}_t^x} \frac{y_j}{|\mathcal{J}_t^x|} \right) \\ &= \frac{1}{|\mathcal{I}|} \sum_{x \in \mathcal{X}} |\mathcal{I}^x| \left( \sum_{i \in \mathcal{I}^x} \frac{\mathsf{y}(i,t)}{|\mathcal{I}^x|} - \sum_{i \in \mathcal{J}_t^x} \frac{y_i}{|\mathcal{J}_t^x|} \right)\\ &\approx_\delta 0. \end{align} Then \begin{align} \frac{1}{|\mathcal{I}|} \sum_{i \in \mathcal{I}} \mathsf{y}(i, t) &\approx_\delta \frac{1}{|\mathcal{I}|} \sum_{i \in \mathcal{I}} p(\mathsf{x}(i),t) \\ &\approx_\epsilon \frac{1}{|\mathcal{J}|} \sum_{i \in \mathcal{J}} p(x_i,t) \\ &= \frac{1}{|\mathcal{J}|} \sum_{x \in \mathcal{X}} |\mathcal{J}^x| \cdot p(x,t)\\ &= \frac{1}{|\mathcal{J}|} \sum_{x \in \mathcal{X}} \frac{|\mathcal{J}^x|}{|\mathcal{J}_t^x|} \sum_{i \in \mathcal{J}_t^x} y_i\\ &= \frac{1}{|\mathcal{J}|} \sum_{i \in \mathcal{J}_t} \frac{y_i}{e_t(x_i)}, \end{align} where ((ref)) uses $\epsilon$-SAP ((ref)).

When the common support assumption ((ref)) is not reasonable, one may be able to use a more coarse-grained approach. Using `coarsened exact matching', we can predict averages not on every $x \in \mathcal{X}$ but on suitable subsets $U \in \Pi$ where $\Pi$ is a partition of $\mathcal{X}$. We discuss this, along with other methods such as doubly robust estimators, instrumental variables, and diff-in-diff, in Appendix (ref).

Comparing the frameworks

In this section, we compare the proposed framework with RCMs. We start with a descriptive comparison in mathematical form, relating the assumptions of $\epsilon$-SAP and $\delta$-CTP with i.i.d. sampling and unconfoundedness (Section (ref)). After that, we give a more subjective account of the practical advantages we see in the new framework (Section (ref)).

Formal considerations

We first note that $\epsilon$-SAP follows from the conventional assumption that covariates in past and future are sampled from the same distribution, at least for bounded $p$ and for large enough populations:

thm\begin{remark}[average prediction in RCMs] \ \\ For datapoints $x_1, ..., x_N$ sampled i.i.d. from a distribution $P$ that governs $X$ on $\mathcal{X}$ and some bounded predictor $p : \mathcal{X} \times \mathcal{T} \to \mathbb{R}$, it follows that $\forall t \in \mathcal{T}$ \begin{equation} \lim_{N \to \infty} \frac{1}{N} \sum_{i=1}^N p(x_i, t) = \mathbb{E}_P[p(X,t)] \end{equation} almost surely. \end{remark}

This means that if both datasets $\mathcal{J}$ and $\mathcal{I}$ are assumed to be sampled i.i.d. from the same distribution over $\mathcal{X}$, this implies that ((ref)) holds almost surely in the limit of infinite data for any $\epsilon>0$.

We can also relate the $\delta$-CTP condition to the RCM framework; this can take different forms, as discussed in the context of different predictors, but often relies on the common unconfoundedness assumption,

equation[equation omitted — 74 chars of source]

It is often suggested that the unconfoundedness setting is “probably the most important one in practice in the modern CI literature” imbens2020. Judea Pearl complains that “I have yet to find a single person who can explain what [it] means in a language spoken by those who need to make this assumption or assess its plausibility in a given problem” pearl2018. Guido imbens2020 cites this passage and shoots back that “simply assuming that one knows or can consistently estimate the joint distribution of all variables in the model” is also “not helpful” (p. 1154). Imbens later adds that unconfoundedness “is so common and well studied that merely referring to its label is probably sufficient for researchers to understand what is being assumed imbens2020. To what extent it is well understood is not easy to settle, but it seems clear that “the unconfoundedness assumption is not directly testable” imbens2024. What our framework provides is an inductive assumption that can be tested in hindsight, which takes different forms for different predictors (see Section (ref) and Appendix (ref)) and can be derived from the high-level idea unconfoundedness:

thm\begin{remark}[conditional means under unconfoundedness] \ \\ For a distribution $P$ over $\mathcal{X} \times \mathcal{Y}_t \times \mathcal{T}$ satisfying unconfoundedness ((ref)) and datapoints $(y_1, t_1), ..., (y_N,t_N)$ sampled i.i.d. from the conditional distribution $P(Y_t,T|x)$, we have, for $\mathcal{J}_t^x(N):=\{1 \leq i \leq N : t_i = t\}$, \begin{equation} \lim_{N \to \infty} \frac{1}{|\mathcal{J}_t^x(N)|} \sum_{i \in \mathcal{J}_t^x(N)} y_i\ =\ \mathbb{E}_P[Y | X=x,T=t] \end{equation} almost surely. This means that $\mu_x(\mathcal{J}, t)$ converges to the conditional expectation for each $x$ and $t$, which is also the limit to which $\mu_x(\mathcal{I}, t)$ converges (trivially). \end{remark}

Thus, for $|\mathcal{J}_t^x|, |\mathcal{I}^x| \to \infty$, $\mu_x(\mathcal{I}, t) - \mu_x(\mathcal{J}, t)$ goes to zero such that ((ref)) is easily satisfied (for any $\delta$) with enough data. Note that this is indeed much stronger than ((ref)) since the latter only requires that the signed average of the differences between the conditional means is small, whereas the remark implies that all absolute differences go to zero. Now ((ref)) also implies that $\delta$-CTP is satisfied for the exact matching predictor, which justifies inverse probability weighting (Proposition (ref)).

While uncounfoundedness is not testable, one can analyse to what extent the validity of results are sensitive to the violation of unconfoundedness in the form of unobserved confounders. In such sensitivity analysis, it is often assumed that outcomes depend on the unobserved confounders via some specified functional form -- particularly common is linearly dependence, since sensitivity analysis is mostly developed for linear regression models. We now show that this is also doable for our framework, and in particular includes the modelling of generalisation to target populations: In the simplest case of assuming linear dependence on unobserved confounders, we may formalise this by adding a term $u \cdot \gamma$ to the $\delta$-CTP:\footnote{This is not too different to the data deficiency coefficient applied to model errors as done in meng2022 (with $r$ seen as a correlation) -- for some further observations on connections to the non-probability sampling literature, see Appendix (ref).}

thm\begin{remark}[sensitivity to unobserved confounders] \ \\ If we relax the $\delta$-CTP assumption by assuming \begin{equation} \mu(\mathcal{I}, t) \approx_\delta \rho(\mathcal{J}, t) + u \cdot \gamma \end{equation} for some unobserved aggregate quantity $u$ that may change between $\mathcal{J}_t$ and $\mathcal{I}$ by some value of $r \in R \subset \mathbb{R}$ and $\gamma \in \Gamma \subset \mathbb{R}$, we can bound the APO by \begin{equation} \mu(\mathcal{I}, t) \geq \rho(\mathcal{J}, t) - (\epsilon + \delta) + \min_{\Gamma, R} \gamma \cdot r, \end{equation} and \begin{equation} \mu(\mathcal{I}, t) \leq \rho(\mathcal{J}, t) + (\epsilon + \delta) + \max_{\Gamma, R} \gamma \cdot r. \end{equation} \end{remark}

Note that this is not too different from the idea of e-values vanderweele2017 for sensitivity analysis, the main difference being our focus on summation rather than ratios; indeed, if we were interested in looser bounds based on maximising over coefficients per-strata $x$ (as for e-values ding2016), we could drop the assumption of a global coefficient $\gamma$ and instead allow a different coefficient for each $x$.

Practical considerations

While sensitivity analyses can be useful, they still do not make the fundamental assumption of (approximate) unconfoundedness testable. Placebo tests go a bit further but rely on further untestable assumptions. A central aspect of the new framework is, then, that inductive assumptions can be formulated in ways that are directly testable and relatable to predictive success -- while still allowing to use intuitions in terms of unconfoundedness to argue for the plausibility of the more concrete assumptions. Another assumption that is easily overlooked is the assumption of (i.i.d.) sampling which, as a relation between observables and a distribution, is difficult to even formalise. While the problem can be seen as less problematic for the pursuit of internal validity (as the distributions is per definition defined for the observed sample), this becomes more problematic when the question of generalisation or transferability to a target population. The new framework avoids this concept altogether (though it can be invoked as a limiting case, as in the related literature on non-probability sampling).

Perhaps most fundamental difference, however, is the shifted aim: from identification to prediction.\footnote{This view has predecessors even for the case of programme evaluation; e.g. berk1987 notes that `evaluations of program impact necessarily involve predictions' which `involves expectations about the likely result under two or more conditions; they are predictions of the what-if variety.’} The notion of identification hinges on the notion of unobserved true parameters. Such parameters do not exist in our framework. This difference also has ramifications for the direct modelling of target populations, testability, the divide between internal and external validity. The lack of testability for assumptions like unconfoundedness has been discussed above. Our proposal is in the spirit of freedman1995 suggesting that `[u]sing models to make predictions of the future, or the results of interventions, would be a valuable corrective', but goes even further by specifying testable assumptions, which makes it possible to distinguish between potential sources of error and making predictions the focus of modelling. This also entails that the target population is explicitly included in the model, which requires (and allows) the modeller to directly grapple with the question of `external validity'.\footnote{We are not claiming that it is impossible in principle to capture generalisation with conventional frameworks. For a review of some recent attempts (and a call for more interest in the problem) see findley2021.}

Another marked difference in our framework is that the separation between internal and external validity dissolves. This is a direct implication of abandoning the idea of identifying supposed true parameters of abstract generative distributions. The concepts of external and internal validity are so engrained in social science research that this may seem extremely counter-intuitive to the experienced researcher. But also this idea is not without precedence: Indeed, the separate consideration of internal and external validity has recently been criticised as a reason for the neglect of external validity (Section (ref)) and the ensuing negative impact on public health research westreich2019. Furthermore, one can recover a version of internal validity by taking the sample population as the target population; we discuss this in the context of difference-in-difference estimators in Appendix (ref).

Another, minor, novel aspect of our framework is that now the APO is the undisputed focus of the statistical enterprise, instead if the ATE. While some estimators explicitly estimate the APOs as a first step, they still tend to be understood as ATE estimators. Our framework more explicitly takes the APOs as the fundamental quantities, the ATE is not more than a formal comparison between APOs.

A more high-level difference is that the new framework explicitly works on averages rather than individuals. In RCMs, ATEs are typically seen as averages of individual treatment effects -- indeed, the subscript $i$ attached to all variables is meant to convey that the parameters and functions directly apply to each individual. Combined with high hopes for Machine Learning techniques (which use a similar formalism), there is an increasing push to `fully personalized treatment effect estimates’ athey2016. This is despite the nature of statistics as a field capturing aggregate behaviour as well as cautionary voices pointing out the complexity of the social world (Section (ref)), where noise cannot be neatly controlled as in physical laboratories berk1987.\footnote{Sander Greenland calls `soft sciences' those that cannot expect to discover numerically precise and general contextual laws analogous to those in physics' greenland2017.} In contrast, our framework dispenses not only with the notion individual effects but also discards the aim of estimating individual quantities; `true' conditional distributions are not defined and the goal is specified as predicting aggregate properties of outcomes in the target population. So far, we have focused exclusively on the APO, i.e. the average, but other aggregate properties are possible, as discussed in the next section.

Beyond APO and ATE

So far, we have only considered APOs, that is, the average outcome if the same treatment is applied to everyone. This is enough to predict the average treatment effect, which is often taken to be the goal of causal inference. There are, however, cases where this is not what we are interested in. We, in turn, discuss relaxations of `same' and of `average'.

Personalised aka conditional treatment rules

Sometimes, we are not interested in applying the same treatment to everyone but instead want make treatment conditional on observed attributes. In this case, we may like to predict the average outcome of more complex covariate-based treatment rules $\pi : \mathcal{X} \to \mathcal{T}$,

equation[equation omitted — 103 chars of source]

In this general formulation, assigning the same treatment rule to everyone, as considered so far, is captured by the degenerate policies $\pi_t : x \mapsto t$ for $t \in \mathcal{T}$. In recent years, starting with manski2004, there has been an increasing focus on learning more sophisticated treatment rules based on data. While we do not discuss the learning part in this paper, we note that our framework straightforwardly applies to predicting the average outcomes of such treatment rules.\footnote{They are sometimes called `individualised' treatment rules -- `conditional' is arguably a better descriptor, as they are conditional on the attributes taken into account (which is a modelling choice), rather than tailored to specific individuals.}

To apply our analysis to such treatment rules $\pi : \mathcal{X} \to \mathcal{T}$, it suffices to consider the sub-populations induced by the level sets of $\pi$, that is,

equation[equation omitted — 94 chars of source]

for $t \in \mathcal{T}$. Based on this, we partition $\mathcal{J}$ and $\mathcal{I}$ into subsets

equation[equation omitted — 215 chars of source]

then apply our machinery to all the pairs $\mathcal{J}^{\mathcal{X}_t}, \mathcal{I}^{\mathcal{X}_t}$ instead of $\mathcal{J}, \mathcal{I}$. For binary treatment $\mathcal{T} = \{0,1\}$, this only means we consider two sub-populations instead of one population. The assumptions connecting $\mathcal{J}^{\mathcal{X}_t}$ and $\mathcal{I}^{\mathcal{X}_t}$ for each $t$ are then analogous to those that we used in this paper to connect $\mathcal{J}$ and $\mathcal{I}$. For example, the assumption that two human populations $\mathcal{J}$ and $\mathcal{I}$ are similar in all relevant respects is hardly weaker than assuming that the respective sub-populations of (for example) people above 40 are similar, as are those below 40. However, this becomes stronger and stronger if we require this for more and more policies $\pi$ simultaneously; requiring this for all possible policies would mean that the two populations must be exactly alike.\footnote{For many other problems such as structural risk minimisation vapnik1982, multi-calibration hebert2018, and randomness vonMises1964, a similar necessity for restricting the number of considered partitions has been observed; see also derr2022.}

Beyond deterministic treatment rules $\pi : \mathcal{X} \to \mathcal{T}$, one may also be interested in more general stochastic treatment rules $\pi : \mathcal{X} \to \Delta(\mathcal{T})$, where $\Delta(\mathcal{T})$ denotes the set of probability distributions over $\mathcal{T}$. For binary treatment decisions, that means that $\pi$ assigns a treatment probability to each $x \in \mathcal{X}$. There are at least two distinct arguments for using such stochastic treatment rules: an ethical, and an epistemic one.\footnote{jain2024 have recently also advocated stochastic allocation rules.} The ethical argument is that hard cut-offs are unfair because two people on opposite sides of the threshold are treated very differently vredenburgh2022. The epistemic argument is that we often do not have much data for some covariates, and if we then never assign some treatments to them, we will never gain that information. This can exacerbate inequality in the case of underrepresented groups, as discussed in oneil2017 and analysed in a bandit setting in li2020. In causal inference, this resurfaces in terms of the assumptions we rely on when analysing data. If we use stochastic treatment rules, we get a well-defined propensity score by design that we can use for subsequent modelling. That is, we satisfy the assumption of `missing at random' rubin1976 (see Footnote (ref)) and get direct access to the conditional probabilities. If the propensity score never reaches 0 or 100 per cent, we also satisfy positivity/common support. We have suggested calibration on sets of equal propensity as a testable criterion. This suggests that it could be beneficial to use treatment rules which only assign a limited set of treatment probabilities, such as $\{0.1, 0.3, 0.5, 0.7, 0.9\}$, instead of the whole interval $[0,1]$, in order to facilitate the analysis.

Beyond averages

In general, a social planner is interested in choosing the policy that maximises desirable properties of the distributions of outcomes. This gives a further reason to focus on treatment-wise potential outcomes rather than supposed treatment effects: Even if meaningful, the distribution of individual treatment effects would not allow us to infer other properties of the outcome distributions beyond the mean manski1996. For simplicity and to directly compare with the bulk of the RCM literature, we have so far restricted ourselves to averages -- but we do consider it important to go beyond this. In the following, we consider three alternatives to the APO, that is, other properties of the potential outcome distribution. This means we consider scalar predictions rather than probabilities for binary outcomes, since for binary outcomes, the average would specify the complete distribution. An example of such an outcome is income.

A first quantity of interest is the proportion of individuals with an income above a certain threshold (such as the poverty line). This just collapses into the APO if we consider the binary outcome of whether an individual has an income above that threshold. Therefore, we do not discuss this property further.

A second potentially interesting property of the outcome distribution is the median or, more generally, quantiles. The target quantity is the $p$-quantile of the (empirical) target distribution for $p \in (0,1)$, that is, the inverse $F_{\mathcal{I},t}^{-1}(p)$ of the cumulative distribution function

equation[equation omitted — 115 chars of source]

This is not necessarily well-defined so we make it more specific by defining

equation[equation omitted — 102 chars of source]

In words, $\xi^t_p$ is the lowest outcome threshold such that at least $p\%$ of the target population fall below it. It turns out that this problem can also be almost reduced to that of APOs -- by fixing the value of the quantile $\xi^t_p$ and then again considering the binary outcome of whether an individual has an income above that threshold. There are two caveats to this reduction. The first one is that we do not see the quantile $\xi^t_p$, so we need to consider the estimator of that quantile as our threshold. Still, by the above reasoning, we can argue that the proportion of people above the chosen threshold should remain similar. The second caveat is that small variation in the proportion may correspond to a large variation in the quantile if the differences between the outcomes around the threshold are high. This requires a further assumption, bound this variation by requiring that for some function $\alpha: \mathbb{R} \to \mathbb{R}$, $\forall y \in \{\mathsf{y}(i,t) : i \in \mathcal{I}\}$,

equation[equation omitted — 152 chars of source]

While it is then analogous to the case of average outcomes (modulo the mentioned changes), we walk through the quantile case in Appendix (ref) as these changes may not be as intuitive.

A third and last type of quantity, that we want to at least hint at, is based on the notion of social welfare functions (SWFs). SWFs take a set of outcomes and reduce them to a number, such as the average, but they may also pay attention to other aspects of the distribution. SWFs that is particularly interesting are rank-dependent and equality-minded kitagawa2021 and can be characterised as

equation[equation omitted — 118 chars of source]

where $Z := \sum_{i=1}^N \omega(F_{\mathcal{I},t}(y_i))$ is a normalising constant and $\omega$ is non-negative and monotonically increasing, i.e. assigning lower weights to higher outcomes based on their rank/percentile. The average corresponds to a constant $\omega$; other SWFs are the minimum, that is, how the worst off are doing, and measures in between. From the more complex characterisation of such outcomes, it is already apparent that it may in general be more difficult to characterise as well as satisfy the required assumptions to predict this with precision. While we consider this an important topic, we leave it for future work.

Discussion

In this paper, we have provided a mathematical framework that takes inspiration from Rubin's potential outcome framework but models the populations without relying on abstract distributions. The focus in this framework shifts from identifying abstract parameters in stipulated distributions to predicting outcomes on the target population. In practice, this allows to directly predict the observable effects of policies and, by relying only on testable assumptions, to retrospectively analyse sources of error in the modelling. These advantages are particularly relevant for evaluating and informing concrete policies and interventions, which are often considered to be the most straightforward application setting for RCMs.

In conformity with Occam's razor, we also avoid unnecessary metaphysical assumptions in the form of well-defined counterfactuals, individual causal effects, or joint probability distributions. We have already discussed the idealising assumption of well-defined probability distributions from which the observations are sampled -- something for which in the complex systems modelled by social and health sciences, `you need a lot of good arguments' cartwright1999. Our framework also suggests, like that of dawid2000, dawid2021, `to reconfigure causal inference as the task of predicting what would happen under a hypothetical future intervention, on the basis of whatever (typically observational) data are available' dawid2022. This is not as radical as it may seem, similar sentiments have been expressed, for example\footnote{Already Ragnar Frisch, `the founding father of modern econometric causal policy analysis' heckman2024 argued that `the scientific [...] problem of causality is essentially a problem regarding our way of thinking, not a problem regarding the nature of the exterior world' frisch1930.}, by berk1987, greenland2012\footnote{Greenland here even believes to discern `a subtle conceptual revolution that recognizes causal inference as a prediction problem' (p. 44).} or hernan2016, who notes that

quote`The goal of the potential outcomes framework is not to identify causes--or to “prove causality”, as it sometimes said. That causality cannot be proven was already forcibly argued by Hume in the 18th century. Rather, quantitative counterfactual inference helps us predict what would happen under different interventions' hernan2016.

These more interpretational considerations should, however, not divert attention from the practical benefits of explicitly modelling the target population and making testable assumption. Indeed, one may see the proposed framework either as a less metaphysically loaded alternative to RCMs or simply as an empirically minded amendment. Although our contribution is only the first step in developing the new version of the framework, we hope to thereby contribute to a solid theoretical basis for using causal inference methods in practice.

Acknowledgements

This work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy—EXC number 2064/1—Project number 390727645 as well as the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A. The authors thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Benedikt Höltgen.

\printbibliography