Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
164,656 characters · 11 sections · 122 citation commands
Experimental Design For Causal Inference Through An Optimization Lens
\CHAPTERNO \TITLE{Experimental Design For Causal Inference Through An Optimization Lens}
\AUBLOCK{ \AUTHOR{Jinglong Zhao} \AFF{Boston University, Questrom School of Business, \EMAIL{[email removed]} } }
\CHAPTERHEAD{Experimental Design for Causal Inference}
\ABSTRACT{ The study of experimental design offers tremendous benefits for answering causal questions across a wide range of applications, including agricultural experiments, clinical trials, industrial experiments, social experiments, and digital experiments. Although valuable in such applications, the costs of experiments often drive experimenters to seek more efficient designs. Recently, experimenters have started to examine such efficiency questions from an optimization perspective, as experimental design problems are fundamentally decision-making problems. This perspective offers a lot of flexibility in leveraging various existing optimization tools to study experimental design problems. This tutorial thus aims to examine the foundations of experimental design problems in the context of causal inference as viewed through an optimization lens. }
\KEYWORDS{ Experimental design, causal inference, optimization }
\vskip 2em
The study of experimental design offers tremendous benefits for answering causal questions across a wide range of applications, including agricultural experiments, clinical trials, industrial experiments, social experiments, and digital experiments. Experimental design is probably one of the cleanest ways to answer causal questions, that is, to understand the causation behind some phenomenon. In such an experiment, the experimenter usually compares the standard offering of some existing policy (the “control”) to a new version of the policy (the “treatment”) by splitting the experimental units into the control and treatment groups. By comparing the outcomes from these two groups, the experimenter discovers the “treatment effect,” that is, the degree to which the newer version is better or worse than the standard version. See Examples (ref) -- (ref) below for an illustration of experimental design terminology in the context of causal inference.
Other than the above examples, experimental design has been recognized as the gold standard in the new product development processes at technology firms koning2022experimentation. Practitioners from a variety of firms have reported their intensive usage of experiments in their product iterations chamandy2016experimentation, chen2024new, farias2023correcting, farronato2018innovation, gupta2019top, huang2023estimating, tang2020control, ye2023deep, zhu2024seller, including search engines (e.g., Bing, Google, Yandex), online retailers (e.g., Amazon, eBay, Etsy), media services (e.g., Netflix), short-form video hosting services (e.g., Douyin, TikTok), social networking services (e.g., Facebook, LinkedIn, Twitter, WeChat), on-demand service platforms (e.g., DoorDash, Lyft, Uber), and travel services (e.g., Airbnb, Booking.com). These firms conduct thousands of new experiments per week to innovate new products and accelerate their product iterations kohavi2013online, tang2010overlapping.
Although experiments have proven beneficial across a wide range of applications, practical constraints such as market size, time, resources, and experimental risks may limit the “sample size” of an experiment --- that is, the total number of experimental units to allocate to the treatment and control groups. Conducting experiments with excessively large sample sizes can be financially and logistically challenging. An emerging question for experimenters, despite their millions and even billions of experimental units, is how to use the experimental units efficiently to draw accurate conclusions.
In recent years, researchers have started to examine such efficiency questions in modern applications through an optimization lens. Experimental design problems are fundamentally decision-making problems bickel2015mathematical, casella2002statistical, neyman1933ix. Given the objectives of the experimenter, the information and uncertainty faced by the experimenter, and the constraints that must be followed, the experimental design problems can be formulated as well-defined optimization problems. Such a perspective offers a lot of flexibility in using various existing optimization tools to study experimental design problems. This manuscript thus aims to examine the foundations of experimental design problems in the context of causal inference as viewed through an optimization lens.
This manuscript is structured as follows. Section (ref) introduces the key elements of experimental design for causal inference and outlines three major optimization frameworks derived from decision theory wald1949statistical, savage1951theory, lindley1956measure, wu1981robustness. We refer to these three frameworks as the robust optimization framework, the stochastic optimization framework, and the deterministic optimization framework. Section (ref) illustrates the robust optimization framework in greater details. Section (ref) introduces additional key elements related to a notion called “covariates,” and uses covariates to illustrate the stochastic optimization framework in greater details. Section (ref) uses covariates to illustrate the deterministic optimization framework in greater details. Section (ref) revisits three basic assumptions from Section (ref); the violation of each assumption leads to many active research directions in modern experimental design for causal inference literature. Section (ref) picks one direction from the violation of each assumption and surveys recent developments in these three directions. We conclude in Section (ref) with recommendations on how to choose the appropriate framework. This manuscript is self-contained, including proofs to all the lemmas and theorems, which can be found in the appendix.
Experimental design is a broad area; as such, this manuscript only covers a narrower scope on experimental design for causal inference. One important omission is the rich literature on “optimal experimental design,” or “optimal design” for short, that draws from the literature of linear algebra, optimization, and statistics. We only discuss the optimal experimental design literature when it overlaps with experimental design for causal inference in Section (ref). We refer to atkinson2007optimum, fedorov2013theory, pukelsheim2006optimal, silvey2013optimal for textbooks and atkinson1975optimal, card1993minimum, titterington1975optimal for papers on optimal experimental design. A related omission is the literature on “optimal Bayesian design.” We only discuss the optimal Bayesian design literature when it overlaps with experimental design for causal inference in Section (ref). We refer to chaloner1982optimal, chaloner1995bayesian, kasy2016experimenters, letham2019constrained, lindley1972bayesian for foundational papers and surveys on optimal Bayesian design.
Causal inference is also a broad area. Here, two major approaches are used to answer causal inference questions: the “potential outcomes” approach and the “structural causal model” approach, and each provide complementary insights. This manuscript only adopts the potential outcomes approach. One important omission is the structural causal model approach. For more details about the two approaches, we refer to ding2023first, hernan2010causal, imbens2015causal, rosenbaum2010design, wager2020stats for textbooks on the potential outcomes approach, and refer to pearl2000models, pearl2016causal, peters2017elements, spirtes2001causation for textbooks on the structural causal model approach.
The mathematical foundation of this manuscript draws from optimization and statistics. For a bigger picture of these two areas, we refer to ben2009robust, bertsekas1997nonlinear, bertsimas1997introduction, birge2011introduction, boyd2004convex, nocedal1999numerical, schrijver2003combinatorial for textbooks on optimization, and berger2013statistical, bickel2015mathematical, buhlmann2011statistics, casella2002statistical, chen2022elements, rigollet2019high, wainwright2019high, wooldridge2010econometric for textbooks on statistics.
Throughout this manuscript, we will use capital letters (e.g., $W$) for random variables and lowercase letters (e.g., $w$) for deterministic quantities. We will use regular font (e.g., $w$) for scalars, bold font (e.g., $\bm{w}$) for vectors, and blackboard font (e.g., $\text{\usefont{U}{bbm}{m}{n}x}$, $\bbbeta$) for matrices. Our notation for matrices is uncommon, yet it is used to distinguish between random variables and deterministic quantities. Unless otherwise stated, we follow the convention that all vectors are column vectors. For any vector $\bm{w}$ or matrix $\text{\usefont{U}{bbm}{m}{n}x}$, we use $\bm{w}^\top$ and $\text{\usefont{U}{bbm}{m}{n}x}^\top$ to stand for their transposes.
We start with one single experimental unit. Let there be two versions of treatments, the active treatment (referred to as “treatment”) and the controlled treatment (referred to as “control”), which we denote using $1$ and $0$, respectively. The random treatment assignment $W$ of this unit takes values from $\{0,1\}$. Following convention, let $W$ stand for a random treatment assignment and $w$ stand for one realization.
One popular framework of causal inference is the potential outcomes framework, which was first proposed in neyman1923application and subsequently studied by holland1986statistics, rubin1974estimating. See athey2017econometrics, ding2023first, imbens2015causal for more bibliographical notes. The potential outcomes framework states that, for this single unit, there exists a pair of two random variables called “potential outcomes,” which each corresponds to a version of treatment. Let $Y(1)$ and $Y(0)$ be the potential outcomes under the treatment assignment and under the control assignment, respectively. Let $(Y(1), Y(0)) \sim \mathcal{F}$ denote that the potential outcomes come from a joint distribution $\mathcal{F}$.
Although there is a pair of two potential outcomes, we can only observe (by drawing a sample from) one of the two potential outcomes. The observed outcome $Y$, sometimes also written as $Y^\mathsf{obs}$, is connected to the potential outcomes by
Note that, whenever we write $Y = Y(W)$ for two random variables, we mean that these two random variables are equal almost surely. The observed outcome $Y$ has two sources of randomness. The first source of randomness comes from the potential outcomes $(Y(1), Y(0))$ as they are sampled from a joint distribution $\mathcal{F}$; the second source of randomness comes from the random treatment assignment $W$.
Under the potential outcomes framework, any comparison of potential outcomes has a causal interpretation. One popular choice of the causal effect, or causal estimand, of interest is the average difference between the potential outcomes under treatment and control,
Because only one of the two potential outcomes can be observed and the other is always missing, the difference $Y(1) - Y(0)$ can never be directly observed. This is also referred to as the fundamental challenge of causal inference by holland1988causal.
Because of this fundamental challenge, the literature usually makes additional assumptions and relies on the availability of multiple units in making causal inference. Now we generalize the above discussion to multiple units. Let there be a total of $n$ units. We refer to $n$ as the “sample size” of an experiment. The experimental units are indexed by $j \in [n] := \{1,2,...,n\}$. For each unit $j \in [n]$, let the random treatment assignment be $W_j$, which takes values from $\{0,1\}$. We collect all treatment assignments in a vector form as $\bm{W}$ that takes values from $\{0,1\}^n$. Following convention, let $\bm{W}$ stand for a vector of random treatment assignments and $\bm{w}$ stand for one realization.
In the most general sense, for each unit $j \in [n]$, there are $2^n$ potential outcomes denoted as $Y_j(\bm{w})$ for all $\bm{w} \in \{0,1\}^n$. We next introduce the non-interference assumption to simplify these many potential outcomes.
The non-interference assumption states that one unit's treatment assignment does not affect the outcomes of any other unit. This assumption dates back to cox1958planning, and serves as a critical component to the stable unit treatment value assumption (SUTVA), which was formally introduced in rubin1980discussion. We refer to ding2023first and imbens2015causal for the full definition of SUTVA. Assumption (ref) holds in many applications such as agricultural experiments, laboratory experiments, and clinical trials, when experiments are perfectly controlled. Assumption (ref) may not always hold in other applications such as social experiments and marketplace experiments. Violation of Assumption (ref) leads to an active research direction on interference. See Section (ref) for further discussions.
We make Assumption (ref) (i.e., the non-interference assumption) throughout this manuscript. Under Assumption (ref), for each unit $j \in [n]$, the observed outcome $Y_j$ is connected to a pair of two potential outcomes $(Y_j(1), Y_j(0))$ by
We collect all $2n$ potential outcomes as $(\bm{Y}(1), \bm{Y}(0))$, which takes values from $\mathcal{Y} \subseteq \mathbb{R}^{2n}$. ding2023first refers to these $2n$ potential outcomes $(\bm{Y}(1), \bm{Y}(0))$ as the “science table,” a term adapted from rubin2005causal.
Although Assumption (ref) greatly simplifies the potential outcomes, we would usually rely on a further assumption on the underlying data-generating process.
Assumption (ref) states that the units are independent and identically distributed. So the observed outcome of one unit is comparable to the observed outcome of any other unit (in the same treatment group). Collecting the observed outcomes from all the units enables us to understand the underlying population. Violation of Assumption (ref) leads to many active research directions in the literature; one of these is treatment heterogeneity. See Section (ref) for further discussions.
Under Assumption (ref) and letting $(Y(1), Y(0)) \sim \mathcal{F}$, we define the average causal effect of the population as
Experimental designs that involve the above causal effect $\tau$, together with Assumption (ref) or some other assumption on the underlying data-generating process, are sometimes referred to as taking a “sampling-based” perspective, so named because the potential outcomes are obtained by sampling from a super-population. Under the sampling-based perspective, we essentially wish to understand the underlying population.
On the other hand, we could also focus on a finite sample and conduct the entire analysis by conditioning on the realized values of the $2n$ potential outcomes; that is, by conditioning on $(\bm{Y}(1), \bm{Y}(0)) = (\bm{y}(1), \bm{y}(0))$. Conditioning on the realized values of the $2n$ potential outcomes, we define the average causal effect of a finite sample as
We make the distinction that whenever we write $\tau_{\bm{Y}(1), \bm{Y}(0)}$, this quantity is random in nature, as the potential outcomes $\bm{Y}(1)$ and $\bm{Y}(0)$ are randomly generated from an underlying data-generating process (Assumption (ref)). But conditioning on one realization $(\bm{Y}(1), \bm{Y}(0)) = (\bm{y}(1), \bm{y}(0))$, the causal effect $\tau_{\bm{y}(1), \bm{y}(0)}$ becomes deterministic. Experimental designs that involve the above causal effect $\tau_{\bm{y}(1), \bm{y}(0)}$ are sometimes referred to as taking a “design-based” perspective, so named because the only source of randomness comes from the design of experiment. Under the design-based perspective, we essentially wish to understand a finite sample.
The sampling-based and the design-based perspectives are highly related. The average causal effect of the population and the average causal effect of a finite sample are connected through the following relationship,
This relationship is because of linearity of expectation and holds even without Assumption (ref). For further discussions of the sampling-based and design-based perspectives, see abadie2020sampling, manski2018right.
Now that we have prescribed the data-generating process, we move on to make assumptions on how to leverage the multiple units. We introduce the following random assignment assumption.
Random assignment is an important property to have. In fact, it is random assignment that warrants the literature of experimental design. One easy way to satisfy the random assignment assumption is through conducting a “randomized experiment.” A randomized experiment $\eta: \{0,1\}^n \rightarrow[0,1]$ induces a joint discrete probability distribution over certain treatment assignment vectors $\bm{w}$, such that
We refer to such a joint discrete probability distribution $\eta$ as a design of experiment. The treatment assignment vector $\bm{W}$ in this experiment conforms to the discrete probability distribution $\eta$; that is, $\Pr(\bm{W} = \bm{w}) = \eta(\bm{w})$.
Assumption (ref) means that $\eta$ must not depend on the potential outcomes of any unit. Violation of Assumption (ref) leads to many active research directions in the literature; one of these is adaptive experiments. See Section (ref) for further discussions.
Several traditional designs of randomized experiments satisfy Assumption (ref), the random assignment assumption. We now introduce two of the simplest designs: the Bernoulli design, and the completely randomized design. Let $\text{\usefont{U}{bbm}{m}{n}1}\{w_j=1\}$ be an indicator function that takes the value of $1$ when $w_j=1$ and takes the value of $0$ otherwise.
Intuitively, in a Bernoulli design, we use independent coin flips to determine the treatment assignments of all the units.
In Definition (ref), each unit has the same treatment probability. As a more general case, each unit could have a different treatment probability that is unrelated to the potential outcomes. For example, in a Bernoulli design, each unit $j \in [n]$ could have a different treatment probability $p_j$, and the joint probability mass function can be written as $\eta(\bm{w}) = \prod_{j=1}^n p_j^{\text{\usefont{U}{bbm}{m}{n}1}\{w_j=1\}} (1-p_j)^{\text{\usefont{U}{bbm}{m}{n}1}\{w_j=0\}}.$
In addition to the Bernoulli design, there is another simple design called the completely randomized design. Let $\dbinom{n}{pn}$ be the binomial coefficient of choosing $pn$ elements from a total of $n$ elements.
Intuitively, a completely randomized design first fixes the number of treatment and control units, and then randomly shuffles the units to determine which units receive treatment.
Different from the Bernoulli design, in a completely randomized design, the treatment assignments are not independent. The treatment assignments of different units are negatively correlated; that is, $\mathrm{Cov}(\text{\usefont{U}{bbm}{m}{n}1}\{W_i=1\}, \text{\usefont{U}{bbm}{m}{n}1}\{W_j=1\}) < 0$ for any $i \ne j \in [n]$. But the marginal distributions of all the treatment assignments are the same; that is, $\Pr(W_i=1) = \Pr(W_j=1)$ for any $i \ne j \in [n]$.
In a randomized experiment, we define, for each unit $j \in [n]$, the “propensity” to be the marginal probability that this unit receives treatment $\Pr(W_j = 1)$. A randomized experiment sometimes satisfies the following assumption on the treatment probability.
We can verify that in the Bernoulli design (Definition (ref)) and completely randomized design (Definition (ref)) above, as long as $p \in (0,1)$, both designs satisfy Assumption (ref).
Once the randomized experiment $\eta$ is determined, we could sample one treatment assignment vector $\bm{w}$ from the distribution induced by $\eta$. Following this treatment assignment vector $\bm{w}$ and after we run the experiment, we could collect the observations and use them to estimate the causal effect. We introduce two of the simplest estimators: the difference-in-means (DM) estimator and the inverse propensity weighting (IPW) estimator.
The difference-in-means estimator, as its name suggests, simply compares the difference between the two sample means of the treatment and control groups. For notational simplicity, we sometimes also write $N(1) = \sum_{j=1}^n \text{\usefont{U}{bbm}{m}{n}1}\{W_j=1\}$ and $N(0) = \sum_{j=1}^n \text{\usefont{U}{bbm}{m}{n}1}\{W_j=0\}$. The difference-in-means estimator $\widehat{\tau}^{DM}$ can then be written as
The IPW estimator, as its name suggests, weighs each observed outcome $Y_j$ by its inverse propensity $\frac{1}{\Pr(W_j=1)}$. When each unit has the same treatment probability, such as in the Bernoulli design (Definition (ref)) and the completely randomized design (Definition (ref)), the IPW estimator $\widehat{\tau}^{IPW}$ can be written as
From this expression we can see that the two estimators only differ by the denominator. The IPW estimator can be seen as replacing the random quantity $\sum_{j=1}^n \text{\usefont{U}{bbm}{m}{n}1}\{W_j=w\}$ in the denominators of the difference-in-means estimator by their expectations $n \Pr\{W_j=w\}$, for $w \in \{0,1\}$, respectively.
For both the Bernoulli design and the completely randomized design, both the difference-in-means estimator and the IPW estimator (leading to a total of four combinations) are well known to be “unbiased,” that is, the expectation of the estimator is equal to the causal effect that we wish to estimate.
So far we have seen that, under each combination of the choice of design of experiment and the choice of estimator, we can always estimate the causal effect (both $\tau_{\bm{Y}(1), \bm{Y}(0)}$ and $\tau$) unbiasedly. A natural question to ask is which combination is more desirable to use. To answer this question from an experimental design perspective, we usually fix the choice of estimator and then choose the optimal design of experiment.
To choose the optimal design of experiment, we adopt the decision-theoretic framework. The decision-theoretic framework was initially introduced by Abraham Wald in the 1940s wolfowitz1952abraham, and was first adopted by wu1981robustness to cast experimental design problems as optimization problems. Under the decision-theoretic framework, we define the following risk function,
where $L(\bm{w}, \bm{y}(1), \bm{y}(0))$ is the loss function. One of the most common choices of loss function is the square loss,
Here we write the estimator in the form of $\widehat{\tau}(\bm{w}, \bm{y}(1), \bm{y}(0))$ to emphasize the dependence on the realized treatment assignments $\bm{w}$ and the realized potential outcomes $\bm{y}(1)$ and $\bm{y}(0)$. Depending on our causal effect of interest, we could also consider a different square loss by replacing the $\tau_{\bm{y}(1), \bm{y}(0)}$ with $\tau$ and taking expectation of the risk function over an additional source of randomness $\mathcal{F}$. The square loss reflects our focus on the quality of estimation. Other common loss functions include penalties associated with false discoveries (a notion related to hypothesis testing), or some costs or losses of revenue if one uses the estimator to directly guide further decisions.
After the loss function is specified, we wish to solve the following minimization problem,
This optimization problem depends on $(\bm{y}(1), \bm{y}(0)) \in \mathbb{R}^{2n}$ (imposing Assumption (ref)) the $2n$ potential outcomes. If one already knew the $2n$ potential outcomes, then one could choose the design optimally by solving the deterministic optimization problem defined above (although if the $2n$ potential outcomes were already known then one could directly report the causal effect without having to run an experiment). But more often, the potential outcomes are unknown. This is essentially a problem of decision making under uncertainty. So we can use a set of different tools to formulate a variety of optimization problems, such as the following two examples,
The first formulation is often referred to as a stochastic optimization problem in the optimization literature, where the underlying data-generating process $\mathcal{F}^n$ is assumed to be known. Here we use $\mathcal{F}^n$ to stand for the joint probability distribution of $\mathcal{F}$ over all $n$ pairs of potential outcomes $(Y_j(1), Y_j(0))$. This formulation is also referred to as Bayes rule in the experimental design literature kasy2016experimenters, lindley1956measure, rubin1978bayesian. The second formulation is often referred to as a robust optimization problem in the optimization literature, where $\mathcal{Y}$, the range of potential outcomes, is assumed to be known. This formulation is also referred to as the minimax rule in the experimental design literature savage1951theory, wu1981robustness.
These different formulations reflect different ways to model the uncertainty governing the potential outcomes $\bm{y}(1)$ and $\bm{y}(0)$. Depending on how we model the uncertainty, we can choose the appropriate framework from (ref) -- (ref). In Sections (ref) -- (ref), we provide examples to illustrate how to model the uncertainty of potential outcomes $\bm{y}(1)$ and $\bm{y}(0)$ in a design of experiment, and cast an experimental design problem as a well-defined optimization problem.
As introduced in Section (ref) formulations (ref) -- (ref), three lines of literature model uncertainty in the potential outcomes from three different perspectives.
The first line of literature often models uncertainty from a robust perspective; that is, there is an adversarial nature that generates the values of the potential outcomes. To restrict this adversarial nature, the potential outcomes will usually take values from a family of candidate values, which are collectively referred to as an “uncertainty set.” The optimal design will usually be randomized in order to achieve minimax optimality. The second line of literature often models uncertainty from a stochastic perspective; that is, the data-generating process of the potential outcomes is assumed to be known, yet the realized values of the potential outcomes are unknown. The optimal design will usually be deterministic as the data-generating process is known. The third line of literature often models uncertainty using specific models; that is, specific models are used to describe the data-generating process, and the uncertainty can be well explained. The optimal design will usually be deterministic as the data-generating process is known and uncertainty can be well explained.
In this section, we illustrate the first line of literature. The other two lines of literature will be illustrated in Sections (ref) and (ref).
This first line of literature dates back to the seminal work of wu1981robustness, who established a deep connection between experimental design and optimization. Below we present a simplified special case of wu1981robustness with only two versions of treatment: one treatment and one control. We adopt the following additive model for the potential outcomes,
where $\alpha_w$ is the effect of treatment $w=1$ or control $w=0$; $g_j$ is the unit fixed effect, which we assume to be deterministic and unknown; and $\epsilon_{jw}$ is the random noise with zero mean and equal variances $\sigma^2$. The assumption that the random noises have equal variances is sometimes also referred to as the “homoscedasticity” assumption. We further assume that the random noises are independent across $j \in [n]$, but are not necessarily independent between $w \in \{0,1\}$. The linear regression literature sometimes specifies a model in the form of $Y_j(w) = \alpha_w + g_j + \epsilon_j$, by implicitly assuming that $\epsilon_{j1} = \epsilon_{j0} = \epsilon_j$. Such a specification is usually equivalent to specifying the additive model (ref), as the specification in (ref) allows the random noises $\epsilon_{j1}$ and $\epsilon_{j0}$ to have arbitrary correlations. More fundamentally, this is because each unit never receives the treatment and control at the same time, so we never interact with $\epsilon_{j1}$ and $\epsilon_{j0}$ at the same time.
Collect $\bm{g} = (g_1, g_2, ..., g_n)$ in a vector form. We assume that $\bm{g} \in \mathcal{G} \subseteq \mathbb{R}^n$ can take any element from a bounded set $\mathcal{G}$. We next make an assumption on $\mathcal{G}$. Denote $\pi: [n] \to [n]$ to be a permutation of $n$ elements; that is, $\pi$ is a one-to-one mapping between its domain and its range. Denote $\Pi$ to be the “permutation group” on $n$ units; that is, it consists of all the $n!$ many possible permutations. Other groups defined on $n$ units, such as a rotation group, are beyond the scope of this manuscript. We refer to good2013permutation for more details. With a little abuse of notation, whenever we apply a permutation $\pi$ to a length-$n$ vector, we permute the elements in the vector; that is, we reload $\pi$ such that $\pi(\bm{g}) = (g_{\pi^{-1}(1)}, g_{\pi^{-1}(2)}, ..., g_{\pi^{-1}(n)})$. Using the above notations, we introduce the following assumption with respect to a permutation group.
Under the model as defined in (ref), we consider the causal effect $\tau = \alpha_1 - \alpha_0$. We also consider the simple difference-in-means estimator $\widehat{\tau} = \widehat{\tau}^{DM} = \frac{1}{N(1)}\sum_{j=1}^n Y_j \text{\usefont{U}{bbm}{m}{n}1}\{W_j=1\} - \frac{1}{N(0)}\sum_{j=1}^n Y_j \text{\usefont{U}{bbm}{m}{n}1}\{W_j=0\}.$ We then consider the loss function, for any realization of the treatment assignment vector $\bm{w}$ and any vector of unit fixed effects $\bm{g}$, to be
Here, because we adopt the model in (ref), we use $\bm{g}$ instead of $\bm{y}(1), \bm{y}(0)$ when writing out the loss function. For any randomized design $\eta: \{0,1\}^n \to [0,1]$, recall the risk function is
We then adopt the robust optimization framework (ref) as introduced at the end of Section (ref) and consider the following minimax optimization problem
To solve the above optimization problem, we make the following observation. For any design $\eta$ and any permutation $\pi$, define $\eta_\pi$ to be a new design such that for any $\bm{w} \in \{0,1\}^n$, $\eta_\pi(\bm{w}) = \eta(\pi(\bm{w}))$. Then, from any design $\eta$, we can construct a new design $\tilde{\eta}$ whose probability mass function is given as follows,
Intuitively, $\tilde{\eta}$ is a design that assigns equal probability to each treatment assignment vector $\bm{w}$ and all its permuted vectors $\pi(\bm{w}), \forall \pi \in \Pi$. More precisely, $\tilde{\eta}$ satisfies the property that $\tilde{\eta}(\bm{w}) = \tilde{\eta}(\pi(\bm{w}))$ for any $\pi \in \Pi$. So $\tilde{\eta}$ can be interpreted as a distribution over completely randomized designs (Definition (ref)). We next show that, under permutation invariance, the optimal design can be represented by a distribution over completely randomized designs.
Now that we establish Lemma (ref), the only flexibility that we have in choosing the optimal design is to decide the sizes of the treatment and control groups. Once the sizes are determined, they naturally determine the completely randomized design. Then, to assign $n$ units to the treatment and control groups, we can choose to balance the sizes of both groups. We formalize the above in the next result.
Permutation invariance is a powerful assumption to establish minimax optimality; bai2023randomize, basse2023minimax have established similar results under permutation invariance. As an alternative, below we introduce a different way of modeling uncertainty, still from a robust optimization perspective. We adopt the robust optimization framework (ref) as introduced at the end of Section (ref) and directly model the potential outcomes $\bm{y}(1), \bm{y}(0)$ to come from a uniform uncertainty set. We will introduce Lemma (ref) and Theorem (ref), which are probably too simple to have been studied in the existing literature. Results in the same spirit have been shown in bojinov2023design, candogan2023correlated, ni2023design yet they do not imply Lemma (ref) and Theorem (ref).
More specifically, we assume the existence of some positive constant $b>0$ such that the potential outcomes are bounded $Y_j(w) \in [-b, b], \forall j \in [n], w \in \{0,1\}$. The range from which the potential outcomes can take values is given by $(\bm{y}(1), \bm{y}(0)) \in \mathcal{Y} = [-b, b]^{2n}$. We consider the causal effect of $\tau_{\bm{y}(1), \bm{y}(0)}$ as defined in (ref). We consider a Bernoulli design parameterized by probabilities $\bm{p} = (p_1, p_2, ..., p_n)$, where each unit $j \in [n]$ received treatment with probability $p_j$ and the treatment assignments across different units are independent. We also consider the IPW estimator $\widehat{\tau} = \widehat{\tau}^{IPW}$ as defined in (ref). We then consider the loss function, for any realization of the treatment assignment vector $\bm{w}$, to be
Note that, in Lemma (ref) and Theorem (ref), we have focused on $\tau$. But here we focus on $\tau_{\bm{y}(1), \bm{y}(0)}$. For any randomized design $\eta_{\bm{p}}: \{0,1\}^n \to [0,1]$, in which we use the subscript to emphasize the dependence on $\bm{p}$, the risk function can be expressed as
We then adopt the robust optimization framework (ref) as introduced at the end of Section (ref) and consider the following minimax optimization problem,
Through expanding the risk function, we can re-write the above optimization problem equivalently as follows.
After this reduction, we can characterize the worst-case potential outcomes and then the vector of optimal treatment probabilities, which determines the optimal Bernoulli design.
Theorems (ref) and (ref) present two examples of casting experimental design problems as robust optimization problems. In both problems, no dominating strategy exists; that is, no design of experiment uniformly achieves the smallest loss over all the potential outcomes. Nonetheless, we can first characterize the worst-case potential outcomes and then find the optimal design of experiment under such worst-case outcomes.
Generally speaking, there is usually no unifying approach in identifying the worst-case potential outcomes or in identifying the optimal design of experiment under the decision-theoretic framework in general. It highly depends on the choice of the loss function, the choice of the causal effect and the estimator, the constraints on the design of experiment, and the framework we adopt in modeling the uncertainty behind the potential outcomes. Theorems (ref) and (ref) illustrate two combinations: Theorem (ref) uses the difference-in-means estimator to estimate the average treatment effect of the population, imposes no constraint on the design of experiment, and assumes permutation invariance; Theorem (ref) uses the IPW estimator to estimate the average treatment effect of a finite sample, considers the family of Bernoulli designs, and assumes a uniform uncertainty set. Other combinations can also be considered, yet the solution approach will likely be different.
The robust optimization framework has many variants. Theorems (ref) and (ref) are two examples of the basic minimax decision rule, which directly minimizes the worst-case risk as in (ref). Other alternatives include the minimax regret decision rule manski2004statistical, stoye2009minimax and the competitive analysis decision rule zhao2023adaptive.
In this section, we have illustrated the robust optimization framework. In Sections (ref) and (ref), we will introduce the other two alternative frameworks of modeling uncertainty, which lead to stochastic and deterministic optimization problems, respectively.
Recall that in Section (ref), we have illustrated the first line of literature that models uncertainty under a robust optimization framework. As a result, the optimal design is usually randomized to achieve minimax optimality. In this section, we illustrate the second line of literature which models uncertainty under a stochastic optimization framework; that is, the data-generating process is usually given and known. The optimal design aims at minimizing the risk function when uncertainty is governed by this data-generating process. In contrast to the robust optimization framework, the optimal design under the stochastic optimization framework is usually deterministic, as the data-generating process is known.
The second line of literature dates back to at least the seminal work of lindley1956measure, in the context of no covariates. To better illustrate the stochastic optimization framework, we formally introduce covariates below. For each unit $j\in[n]$, let there be an associated vector of covariates $\bm{X}_j$ that takes values from $\mathcal{X} \subseteq \mathbb{R}^d$. We refer to $d$ as the “dimension” of the covariates. Throughout this manuscript, we assume that dimension $d$ is much smaller than sample size $n$. Such problems are often referred to as low-dimensional problems. When dimension $d$ is comparable to, or even larger than, sample size $n$, it leads to high-dimensional problems. We refer to buhlmann2011statistics, rigollet2019high, tibshirani1996regression, wainwright2019high for further discussions.
We start by introducing an assumption on the underlying data-generating process of the covariates. Similar to Assumption (ref), we assume that the covariates of different units are also generated from the same distribution.
\begin{assumption+}{(ref)$^*$}[Homogeneity] For each unit $j \in [n]$, the potential outcomes and the covariates $(Y_j(1), Y_j(0), \bm{X}_j^\top)$ are independent and identical samples drawn from the same super-population, that is,
\end{assumption+}
To show the value of modeling covariates, we start from the following example in wager2020stats about cash incentives for non-smoking.
If we carefully examine Example (ref), we will see that Assumption (ref) (i.e., the random assignment assumption) fails in the context of Simpson's paradox. In the combined data, the potential outcomes are correlated with the treatment assignments because the treatment probabilities are different across the two locations, and the two locations also have different baseline smoking rates. In this example, the location is referred to as a “confounder.”
The potential existence of confounders motivates the conditional random assignment or the unconfoundedness assumption.
\begin{assumption+}{(ref)$^*$}[Unconfoundedness] For each unit $j \in [n]$, conditional on the covariates $\bm{X}_j$, the treatment assignment $W_j$ and the potential outcomes of all units $\big\{(Y_j(1), Y_j(0))\big\}_{j=1}^n$ are independent, that is,
\end{assumption+}
In Example (ref), let $X_j$ be one single covariate indicating whether the experiment is conducted in Geneva or in Palo Alto. Then, conditional on the location $X_j$, the treatment assignment is as good as random. More generally, if the random treatment assignment only depends on covariates, we can generalize the definition of the propensity score from Section (ref) and define the propensity score as a function of covariates. Define, for any unit $j \in [n]$ and conditional on its covariates taking values $\bm{X}_j = \bm{x}$, the propensity score $e(\bm{x})$ to be the probability that this unit receives treatment; that is, $e(\bm{x}) = \Pr(W_j = 1 \vert \bm{X}_j = \bm{x})$. A randomized experiment sometimes satisfies the following assumption on the propensity score.
\begin{assumption+}{(ref)$^*$}[Positivity] For any $\bm{x} \in \mathcal{X} \subseteq \mathbb{R}^d$, the propensity score is strictly between $0$ and $1$, that is,
\end{assumption+}
Simpson's Paradox in Example (ref) and the definition of the propensity score motivate a design of experiment that depends on the covariates, which we refer to as a “stratified randomized design.”
Each stratified randomized experiment is parameterized by a $k$-partition $\mathcal{P}$ and a $k$-dimensional vector $\bm{p} \in (0,1)^k$. In many applications, and similar to Example (ref), the treatment probabilities are different across different strata. When the treatment probabilities are different across different strata, the traditional difference-in-means estimator $\widehat{\tau}^{DM}$ as defined in Definition (ref) is no longer unbiased. As an alternative, we aggregate the difference-in-means estimators from each stratum.
We can verify that, conditional on each stratum having at least one treatment unit and one control unit, the aggregate estimator $\widehat{\tau}^{AGG}$ is unbiased. On the other hand, as long as Assumption (ref) holds, the IPW estimator $\widehat{\tau}^{IPW}$ is also unbiased. Both claims are similar to Lemma (ref) and their proofs follow similarly.
In some other applications, the treatment probabilities are the same across different strata; that is, $p_l = p, \forall l \in [k]$. When the treatment probabilities are the same, the traditional difference-in-means estimator $\widehat{\tau}^{DM}$ is unbiased. One simple special case is when we set $p_l = \frac{1}{2}, \ \forall l \in [k]$. Once $p_l = \frac{1}{2}, \ \forall l \in [k]$ is fixed, the key decision is then to choose the partition $\mathcal{P}$.
Choosing $\mathcal{P}$ can be done in different ways. Next, we show how to choose $\mathcal{P}$ from an optimization framework following bai2022optimality. A similar idea also appeared in kallus2018optimal. We make Assumption (ref), that is, the homogeneity assumption. We then observe the covariates $\bm{x}_1, ..., \bm{x}_n$ and collect them into matrix $\text{\usefont{U}{bbm}{m}{n}x} \in \mathbb{R}^{n \times d}$ by stacking $\bm{x}_1^\top, ..., \bm{x}_n^\top$ by rows. We would like to partition the units into different strata after observing their covariates. We emphasize that the partition $\mathcal{P}$ is after we observe all the covariates $\text{\usefont{U}{bbm}{m}{n}x}$, and the optimal partition $\mathcal{P}$ should depend on the observed covariates $\text{\usefont{U}{bbm}{m}{n}x}$. Yet we drop the dependence on $\text{\usefont{U}{bbm}{m}{n}x}$ from $\mathcal{P}$ when writing the partition. We assume that the treatment probabilities are $\frac{1}{2}$ across all strata, and within each stratum we conduct a completely randomized experiment that is independent of any other stratum.
We consider the following average causal effect conditional on the covariates from a finite sample
conditional on all the covariates $\text{\usefont{U}{bbm}{m}{n}x}$. Here we use $\mathcal{F} \vert \bm{X}_j$ to emphasize that conditional on the covariates being equal to $\bm{x}_j$, the potential outcomes of unit $j \in [n]$ may have a different distribution than unconditionally drawn from $\mathcal{F}$. So $\tau_{\text{\usefont{U}{bbm}{m}{n}x}}$ may be different from $\tau = \mathrm{E}[Y(1) - Y(0)]$ the causal effect of the underlying population. But if we further take expectation then $\mathrm{E}_{\text{\usefont{U}{bbm}{m}{n}x} \sim \mathcal{F}}[\tau_{\text{\usefont{U}{bbm}{m}{n}x}}] = \tau$.
We then choose the estimator to estimate the causal effect $\tau_{\text{\usefont{U}{bbm}{m}{n}x}}$. Given the treatment probabilities are all equal to $\frac{1}{2}$ from all strata (so the difference-in-means estimator is unbiased), we consider the difference-in-means estimator $\widehat{\tau} = \widehat{\tau}^{DM} = \frac{1}{N(1)}\sum_{j=1}^n Y_j \text{\usefont{U}{bbm}{m}{n}1}\{W_j=1\} - \frac{1}{N(0)}\sum_{j=1}^n Y_j \text{\usefont{U}{bbm}{m}{n}1}\{W_j=0\}.$
We then consider the loss function, for any realization of the treatment assignment vector $\bm{w}$ and any realization of the potential outcome vectors $\bm{y}(1)$ and $\bm{y}(0)$, to be the square loss,
For any stratified randomized design $\eta_{\mathcal{P}}$, in which we use $\mathcal{P}$ in the subscript to emphasize the dependence on the strata, and any realization of the potential outcome vectors $\bm{y}(1)$ and $\bm{y}(0)$, we write the risk function as
Finally, we adopt the stochastic optimization framework (ref) as introduced at the end of Section (ref) and consider the following stochastic optimization problem,
where we use $\mathcal{F}^n \vert \text{\usefont{U}{bbm}{m}{n}X}$ to stand for the joint probability distribution of $\mathcal{F}$ over all the $n$ units, conditional on their covariates $\text{\usefont{U}{bbm}{m}{n}X}$. We use $\mathrm{E}_{\eta_{\mathcal{P}}, \mathcal{F}^n \vert \text{\usefont{U}{bbm}{m}{n}X}}\big[ (\widehat{\tau} - \tau_{\text{\usefont{U}{bbm}{m}{n}x}})^2 \vert \text{\usefont{U}{bbm}{m}{n}x} \big]$ to emphasize that randomness has two sources, one coming from the stratified randomization, and the other coming from the potential outcomes. Note that the randomness of potential outcomes is conditional on the covariates taking values $\text{\usefont{U}{bbm}{m}{n}X} = \text{\usefont{U}{bbm}{m}{n}x}$. The decision that we make, $\mathcal{P}$, should also depend on $\text{\usefont{U}{bbm}{m}{n}x}$, but we make it implicit.
In what follows, we expand on the mean squared error as defined above in (ref). Before we proceed, we examine a simple result that holds generally.
Using Lemma (ref), the mean squared error of any estimator can be decomposed into two terms, its variance term, and the square of its bias term. This decomposition usually allows for the study of the bias-variance trade-off.
Using results similar to Lemma (ref), we show that minimizing the mean squared error in the objective function of (ref) is equivalent to the following minimization problem.
Now we define
to be the baseline function. In Example (ref), this function stands for the baseline smoke rate in Geneva or Palo Alto. For each unit $j \in [n]$, let the baseline be $g_j = g(\bm{x}_j)$ and collect all the baselines in a vector form as $\bm{g} = (g_1, g_2, ..., g_n)$. If the baselines are known and given, we can use the baselines to re-write the expression in Lemma (ref) as
where $\text{\usefont{U}{bbm}{m}{n}v}$ is the following covariance matrix
Because the randomization is independent across different strata, the covariance between the treatment assignments of any two units across two different strata is zero. So the covariance matrix $\text{\usefont{U}{bbm}{m}{n}v}$ can be decomposed into several diagonal blocks, with the off-diagonal components being equal to zero. We can then look into the diagonal blocks and show that the optimal partition consists of only size-two strata.
Theorem (ref) presents an example of how to cast a covariate-dependent experimental design problem as a stochastic optimization problem. Usually, the optimal solution to a stochastic optimization problem is deterministic unless the problem has multiple optimal solutions. As we have seen from Theorem (ref), because our decision space is over the stratification, the optimal stratification is deterministic. However, in a stratified randomized experiment, we conduct a randomized experiment within each stratum. Even though the optimal decision (i.e., stratification) is deterministic, the experiment itself is still randomized as a result of randomization within each stratum.
The optimal matched pair design as in Theorem (ref) is a special case of the stochastic optimization framework. The optimal design only requires knowledge about the baseline function $g(\cdot)$, which only depends on the first moments of the joint distribution. The optimal design does not require knowledge about the entire joint distribution $(Y_j(1), Y_j(0)) \sim \mathcal{F} \vert \bm{X}_j$. Generally speaking, however, the joint distribution will probably be required to find the optimal design under the stochastic optimization framework.
To conclude this section, we point out that we have used the baseline function $g(\cdot)$ to guide the randomization. In Section (ref) we will take a deeper look at the baseline function $g(\cdot)$ by imposing a linear model, which sometimes leads to deterministic optimization problems.
Recall that in Section (ref), we have illustrated the first line of literature that models uncertainty under a robust optimization framework; in Section (ref), we have illustrated the second line of literature that models uncertainty under a stochastic optimization framework. In this section, we illustrate the third line of literature, which uses specific models to explain uncertainty. As uncertainty can be well explained, the optimal design is usually deterministic.
The third line of literature dates back to the seminal work of smith1918standard. Following the notations in sibson1974optimality, we adopt a linear additive model that is similar to (ref) as described in Section (ref) to incorporate covariates,
where $\tau$ is the causal effect that we are interested in, which corresponds to $(\alpha_1 - \alpha_0)$ in model (ref); $\bm{X}_j$ is a vector of covariates of unit $j$; $\bm{\beta}$ is a vector of unknown parameters; and $\epsilon_{jw}$ is the random noise with zero mean and equal variance $\sigma^2$. The random noises are independent across $j \in [n]$, but are not necessarily independent between $w \in \{0,1\}$. We do not specify a fixed effect here, because for any unit $j\in[n]$, we can always augment the vector $\bm{X}_j$ by adding one extra dimension to $\bm{X}_j$ such that the first dimension of $\bm{X}_j$ is equal to $1$. Such an augmentation allows $\bm{\beta}^\top\bm{X}_j$ to incorporate a non-zero intercept.
Under the linear additive model (ref), we usually use linear regression to estimate the causal effect $\tau$ and the coefficients $\bm{\beta}$. To succinctly describe the linear regression, we introduce the following notations. We collect causal effect $\tau$ and coefficients $\bm{\beta}$ into a vector and denote $\bm{\theta} = (\tau, \bm{\beta}^\top)^\top \in \mathbb{R}^{d+1}$. We collect all covariates into a matrix $\text{\usefont{U}{bbm}{m}{n}X} \in \mathbb{R}^{n \times d}$ by stacking $\bm{X}_1^\top, ..., \bm{X}_n^\top$ by rows. For any random treatment assignment vector $\bm{W}$ (which possibly depends on the covariates $\text{\usefont{U}{bbm}{m}{n}X}$), we collect them with $\text{\usefont{U}{bbm}{m}{n}X}$ and denote $\text{\usefont{U}{bbm}{m}{n}Z} = \big[ \bm{W} \ \text{\usefont{U}{bbm}{m}{n}X} \big]$, which takes values from $\{0,1\}^{n \times 1} \times \mathbb{R}^{n \times d}$. We also collect all observed outcomes $Y_j, \forall j \in [n]$ into a vector $\bm{Y} \in \mathbb{R}^n$. Let $\epsilon_{j1} = \epsilon_{j0}, \forall j \in [n]$ and denote $\epsilon_j = \epsilon_{j1} = \epsilon_{j0}, \forall j \in [n]$. We collect all random noises $\epsilon_{j}, \forall j \in [n]$ into a vector $\bm{\epsilon}$ that takes values from $\mathbb{R}^n$. We then have (ref) on the $n$ samples,
Under the homoscedasticity assumption, we usually estimate $\bm{\theta}$ using the ordinary least squares (OLS) estimator, which is defined as follows.
It is well-known that the OLS estimator is unbiased. To see this, we re-write the OLS estimator as
Because the noises $\epsilon_{j}, \forall j \in [n]$ all have zero means, the OLS estimator is unbiased, that is,
The source of randomness behind this unbiasedness comes from the random noises $\bm{\epsilon}$. Regardless of the design of experiment (as long as $\text{\usefont{U}{bbm}{m}{n}Z}^\top \text{\usefont{U}{bbm}{m}{n}Z}$ is invertible), this unbiasedness result holds.
Below we study how to choose a design of experiment when we use the OLS estimator. Suppose we are given $\bm{x}_1, ..., \bm{x}_n$, the realizations of the covariates from the $n$ units. We now wish to assign all $n$ units into treatment and control groups. To do so, we start with one simple objective.
Under the model as defined in (ref), we consider the causal effect $\tau$. We also consider the OLS estimator $\widehat{\tau} = \widehat{\tau}^{OLS}$, which is the first dimension of the estimator $\widehat{\bm{\theta}}$. We then consider the loss function, for any realization of the treatment assignment vector $\bm{w}$, to be
The second equality holds because of expression (ref); that is, the OLS estimator is unbiased. Note that as the loss function is deterministic, the optimal design $\eta: \{0,1\}^n \to [0,1]$ must also be deterministic; that is, the optimal design must put all probability mass on one support unless the problem has multiple optimal solutions. So the risk minimization problem is equivalent to the following deterministic optimization problem,
Problem (ref) above is sometimes referred to as the $D_A$-optimal experimental design problem.
To solve the $D_A$-optimal experimental design problem, we examine the covariance matrix of the OLS estimator. Below we write $\text{\usefont{U}{bbm}{m}{n}z}$ and $\text{\usefont{U}{bbm}{m}{n}x}$ to emphasize that we are given the realizations of the covariates $\bm{x}_1,...,\bm{x}_n$. Note that
where the second equality is because $\mathrm{Var}_{\bm{\epsilon}}(\bm{\epsilon}\bm{\epsilon}^\top)$ is an $n \times n$ diagonal matrix with diagonal elements equal to $\sigma^2$ (assuming homoscedasticity) and the off-diagonal elements equal to $0$. The variance of $\widehat{\tau}$ (the first element of $\widehat{\bm{\theta}}$) is then equal to the first element in the first row of the matrix $\mathrm{Var}_{\bm{\epsilon}}\big(\widehat{\bm{\theta}}\big)$, that is, $\sigma^2 \bm{e}_1^\top (\text{\usefont{U}{bbm}{m}{n}z}^\top \text{\usefont{U}{bbm}{m}{n}z})^{-1} \bm{e}_1$, where $\bm{e}_1 = (1, 0, 0, ..., 0)^\top$ stands for a basis vector with only a one as the first element and zero otherwise. Using the bordering method in block matrix inversion, we can calculate $(\text{\usefont{U}{bbm}{m}{n}z}^\top \text{\usefont{U}{bbm}{m}{n}z})^{-1}$ and obtain
Many algorithms have been used to (approximately) solve the $D_A$-optimal design problem. One algorithm, given by bhat2020near, is to reduce this problem to the maximum cut problem, which converts the goemans1995improved result into a $\frac{2}{\pi}$-approximate solution to this $D_A$-optimal design problem. We skip the details of the approximation algorithm as they go beyond the scope of this manuscript.
The optimal experimental design problem has a different version, and this different version is probably more widely studied in the literature. In the above derivation, we allocate all the $n$ units into either the treatment or the control group. We then use all the data from these $n$ units to construct the OLS estimator to estimate the causal effect $\tau$. In the different version of the optimal experimental design problem, the decision is to select a subset of the $n$ units to conduct experiments, and only one version of treatment (rather than treatment and control) is involved in the experiment. Because there is only one version of treatment, the causal effect $\tau$ is not well-defined, and the only focus is on estimating the unknown parameters $\bm{\beta}$.
Mathematically, letting $\bm{x}_j$ be the covariates for unit $j \in [n]$, the decision is to choose a subset $S \subseteq [n]$, with at most cardinality $|S| \leq k$, to construct the OLS estimator,
The objective is related to the covariance matrix of $\mathrm{Var}_{\bm{\epsilon}}\big(\widehat{\bm{\beta}}\big)$. The $D_A$-optimal design problem is concerned with minimizing a scalarization of the covariance matrix $\mathrm{Var}_{\bm{\epsilon}}(\widehat{\bm{\beta}})$ on some dimension $l \in [d]$. Many other objectives can be considered. Among them, one of the most popular is the $D$-optimal criterion fedorov2013theory, pukelsheim2006optimal, silvey2013optimal. Recall from (ref) that
The $D$-optimal design minimizes the determinant of this covariance matrix, where the letter $D$ stands for “determinant.” The $D$-optimal design also has a geometric interpretation of minimizing the volume of an ellipsoid at any fixed confidence level titterington1975optimal. Using the fact that the determinant of the inverse of an invertible matrix is the reciprocal of the determinant of the matrix, that is, $\det(\text{\usefont{U}{bbm}{m}{n}A}^{-1}) = \det(\text{\usefont{U}{bbm}{m}{n}A})^{-1}$, the $D$-optimal design is equivalent to minimizing
This optimization problem is computationally challenging, and many efforts have been made to computationally solve this problem. We refer to allen2021near, madan2019combinatorial, meyer1995coordinate, nikolov2015randomized, nikolov2016maximizing, singh2018approximate, singh2020approximation, summa2014largest for recent developments.
So far we have seen the covariates that are directly observed. At times, we may also worry about the unobserved covariates. One way to incorporate the unobserved covariates is through historical data, using the synthetic control method abadie2021using, abadie2010synthetic, abadie2003synthetic. To introduce the synthetic control method, we first introduce the panel data. Let there be $n$ units, denoted as $[n]$. We focus on a setting where these $n$ units are fixed (instead of randomly sampled from a super-population). Let there be $T$ discrete, finite time periods, denoted as $[T]$. In the panel data setting, the convention is to use the capital letter $T$ for a fixed time horizon; even though we use the capital letter $T$, it is still a constant, not a random variable. In panel data, each unit $j \in [n]$ is repeatedly exposed to treatment or control, as well as repeatedly observed for a duration of time. For each unit $j \in [n]$ and at each time period $t \in [T]$, we denote the treatment assignment as $W_{jt}$, which takes values from $\{0,1\}$, and the observed outcome as $Y_{jt}$, which takes values from $\mathbb{R}$. The observed outcomes are connected to the potential outcomes by
Suppose we stand at the end of time period $T_0$. We refer to such $T_0$ periods as the pre-experimental periods. The last $(T-T_0)$ periods $\{T_0+1, ..., T\}$ are the experimental periods. All units receive control during the pre-experimental periods; that is, $W_{jt} = 0, \forall j \in [n], t \in [T_0]$. At the end of time period $T_0$ and after collecting the observed outcomes up until period $T_0$, we choose a subset of the $n$ units to receive treatment during the experimental periods, leaving the remaining units in control.
We adopt a linear model to incorporate both observed covariates and unobserved covariates. Such a model is often referred to as a linear factor model.
where $\alpha_t(w)$ stands for a fixed effect; $\bm{X}_j$ stands for a column vector of observed covariates that takes values from $\mathbb{R}^d$; $\bm{\mu}_j$ stands for a column vector of unobserved covariates that takes values from $\mathbb{R}^r$; $\bm{\beta}_t(w) \in \mathbb{R}^d$ and $\bm{\lambda}_t(w) \in \mathbb{R}^r$ stand for two column vectors of unknown parameters; and $\epsilon_{jt}(w)$ is the random noise with zero mean and sub-Gaussian tail whose variance proxy is upper bounded by $\overline{\sigma}^2$. The random noises are independent across $j \in [n]$, but are not necessarily independent between $w \in \{0,1\}$.
We use the synthetic control method to choose the treatment and control units, and to estimate the causal effect. For each unit $j \in [n]$ and at any time period $t \in [T]$, let $\tau_{jt} = Y_{jt}(1) - Y_{jt}(0)$ be the individual causal effect. We are given a vector of constants $\bm{f} = (f_1, ..., f_n) \in [0,1]^n$, such that $\sum_{j=1}^n f_j = 1$. We are interested in the following average treatment effect for any $t \in \{T_0+1,...,T\}$,
In this setting, the synthetic control estimator takes as inputs two vectors of weights, $\bm{u} = (u_1, ..., u_n), \bm{v} = (v_1, ..., v_n)\in [0,1]^n$, such that $\sum_{j=1}^n u_j = \sum_{j=1}^n v_j = 1$ and $u_j v_j = 0, \forall j \in [n]$. Once $\bm{u}$ and $\bm{v}$ are determined, the synthetic control estimator for any $t \in \{T_0+1,...,T\}$ is given by
Intuitively, the synthetic control estimator aims at constructing a synthetic treatment unit $\sum_{j=1}^n u_j Y_{jt}$ that mimics the unobservable averaged treatment outcome $\sum_{j=1}^n f_j Y_{jt}(1)$, as well as a synthetic control unit $\sum_{j=1}^n v_j Y_{jt}$ that mimics the unobservable averaged control outcome $\sum_{j=1}^n f_j Y_{jt}(0)$. In observational studies, we usually only construct the synthetic control unit, and this is where the name “synthetic control” comes from.
We next introduce a quadratic program to determine the weights $\bm{u}$ and $\bm{v}$. Let $k \leq n$ be a small positive integer that restricts the number of treatment units. For any $j \in [n]$, let $\bm{Y}_j = (Y_{j1}, Y_{j2}, ..., Y_{jT_0})$ be the vector of observed outcomes during the pre-experimental periods. We then denote $\bm{Z}_j = (\bm{Y}_j^\top, \bm{X}_j^\top)^\top$ and $\overline{\bm{Z}} = \sum_{j=1}^n f_j \bm{Z}_j$. At the end of period $T_0$, conditional on $\bm{Y}_j = \bm{y}_j$ the observed outcomes during the pre-experimental periods and $\bm{X}_j = \bm{x}_j$ the realizations of the covariates, we write the following quadratic program. The decision variables to the quadratic program are $\bm{u}$ and $\bm{v}$, the weights of the synthetic control estimator.
After solving (ref), we conduct an experiment and assign units that have a positive $u_j > 0$ weight into the treatment group and units that have a zero $u_j = 0$ weight into the control group. Out of those units in the control group, we use the units that have a positive $v_j > 0$ weight to form the synthetic control unit. We refer to the above design of experiment as the synthetic control design. The synthetic control design as shown in (ref), as well as other formulations of the synthetic control design, such as the ones given in doudchenko2019designing, doudchenko2021synthetic, is computationally challenging. Recently, efforts have been made to computationally solve such problems lu2022synthetic.
Denote the optimal solution to (ref) as $(\bm{u}^*, \bm{v}^*)$. The quality of the optimal solution $(\bm{u}^*, \bm{v}^*)$ will have an impact on the performance of the estimator $\widehat{\tau}^{SC}_t(\bm{u}^*, \bm{v}^*)$. Before formally quantifying the performance of the estimator, we introduce a few more notations. Denote $\overline{\beta} = \max_{j \in [n], t \in [T], w \in \{0,1\}} \beta_{jt}(w)$ and $\overline{\lambda} = \max_{j \in [n], t \in [T], w \in \{0,1\}} \lambda_{jt}(w)$. Denote $\bblambda(0)$ to be a $(T_0 \times r)$ matrix whose $t$-th row is equal to $\bm{\lambda}_t(0)^\top$. Denote $\zeta_{T_0}$ to be the smallest eigenvalue of $\bblambda(0)^\top \bblambda(0)$.
Theorem (ref) suggests that, the ex-post bias of the synthetic control estimator $\widehat{\tau}^{SC}_t(\bm{u}^*, \bm{v}^*)$ consists of two parts. The first part depends linearly on $c$, which reflects the quality of $(\bm{u}^*, \bm{v}^*)$ in minimizing the quadratic program (ref). The second part depends on the random noises, and it decreases as $T_0$, the number of pre-experimental periods, increases.
To conclude this section, we note that Theorem (ref) only presents an upper bound of the ex-post bias of the synthetic control estimator; it does not fully characterize the true ex-post bias. Yet it is probably the best characterization in the synthetic control literature. Sometimes when the performance of the estimator is challenging to directly optimize, we optimize its proxies instead.
The classical causal inference literature makes Assumptions (ref) -- (ref). So far, we have seen different optimization frameworks under these three classical assumptions. These assumptions may not always hold in modern applications. Below we survey recent developments in experimental design for causal inference when these classical assumptions are violated.
Violation of Assumption (ref) leads to a rich literature on interference hudgens2008toward, tchetgen2012causal. Intuitively, Assumption (ref) fails because the treatment assignment of one unit may have an impact on the potential outcomes of other units. In the full generality, it requires all $n$ treatment assignments to describe the potential outcomes of each unit.
Depending on the causes, interference may be modeled in different ways. Probably the most popular way of modeling interference is through a network, which leads to the network interference literature. The network interference literature models each unit as a vertex on a network. The treatment assignment on one unit may have an impact on the potential outcomes of other units through the edges of the network, which is a phenomenon called “spillover” effects. One popular approach of studying network interference is through a concept called “exposure mapping” introduced by aronow2017estimating, which develops the “constant treatment response” assumption introduced by manski2013identification. Exposure mapping is a dimension reduction mapping that reduces the dependence from all $n$ treatment assignments to a much smaller number of quantities. See Example (ref) for an illustration of exposure mapping.
Example (ref) only illustrates one specification of exposure mapping. Exposure mapping can be specified in various other ways. When the exposure mapping is well-specified and known, the network interference literature has examined different designs of experiments. Most of these experiments share one similar idea and can be seen as some variants of “cluster randomized experiments.” In a cluster randomized experiment, all the units (vertices) are first grouped into multiple clusters. Then, each cluster is randomly assigned into treatment or control, so that all the units within the same cluster receive the same version of treatment assignments. See brennan2022cluster, candogan2023correlated, cortez2023exploiting, eckles2016estimating, eckles2017design, eichhorn2024low, jagadeesan2020designs, han2024population, harshaw2023design, holtz2020limiting, holtz2024reducing, jiang2023causal, leung2022rate, leung2023design, ni2023design, pouget2018optimizing, pouget2019variance, qu2021efficient, rolnick2019randomized, saveski2017detecting, ugander2013graph, ugander2023randomized, viviano2020experimental, viviano2023causal, and references therein.
Sometimes the exposure mapping is unknown or potentially misspecified. Recent works have examined various estimation strategies without using knowledge of the exposure mapping. Yet, they need to assume that spillover effects decay with respect to the distance on the network; that is, spillover effects are only local. See belloni2022neighborhood, leung2022causal, leung2022rate, savje2021average, yu2022estimating, yuan2021causal, yuan2023two. Because exposure mapping is unknown, the works in this literature usually conduct simple experiments and combine them with non-trivial analysis, with the exception of leung2022rate, who studies how to choose the rate-optimal design of experiment. If one can conduct multiple experiments on the same network over time, recent works consider rolling out experiments over time to make estimation and inference in the presence of unknown interference. See boyarsky2023modeling, cortez2024combining, han2022detecting.
Other than network interference, the interference literature has gained increasing popularity in modern marketplace applications, which leads to the literature on marketplace interference. Marketplace interference can sometimes be modeled as network interference. For example, the market equilibrium prices or the distributional interactions (Example (ref)) can sometimes be used as exposure mappings. See bajari2021multiple, munro2021treatment, munro2024treatment, wager2021experimenting. On the other hand, marketplace interference can sometimes be captured by explicit models, such as auction models, discrete choice models, game-theoretic models, market-making models, and queuing models. See basse2016randomization, bright2022reducing, dhaouadi2023price, johari2022experimental, kuang2024detecting, liao2023statistical, li2022interference, li2023experimenting, si2023optimal, ye2023cold.
Recently, the interference literature also gains increasing popularity in recommendation systems, which leads to the literature on feedback loops, or symbiosis biases. See goli2023bias, holtz2023study, johnson2017ghost, si2023tackling, zhan2024estimating. In a recommendation system, for example, user data obtained under previous recommendations are used as inputs to re-train the recommendation system. Unless treatment and control data are separately used to train two systems, interference occurs from sharing a common pool of training data. See si2023tackling for an excellent review of feedback loops and related interference literature.
One remarkable special case of interference is the literature on temporal experiments, in which we can find temporal analogues of network interference and marketplace interference. The temporal analogue of network interference is switchback experiments, which have appeared under various different names such as n-of-1 trials liang2023randomization, time series experiments bojinov2019time, and crossover designs basse2023minimax. A switchback experiment is a special case of network interference because the graph of interference in a switchback experiment follows a single line. The spillover effects are referred to as “carryover effects” in this setting. In earlier works, switchback experiments are used in agricultural applications to compare the effects of different feeding plans on milk yields cochran1941double. Recently, switchback experiments have gained increasing popularity as a result of the rise of modern applications such as on-demand service platforms (e.g., DoorDash, Lyft, and Uber; see chamandy2016experimentation, kastelman2018switchback for blog posts on this topic). glynn2020adaptive is the first to study the design of switchback experiments in such modern applications, followed by bojinov2023design, chen2023switchback, hu2022switchback, jia2023faster, xiong2024data. See hu2022switchback for an excellent introduction of switchback experiments.
The temporal analogue of marketplace interference involves modeling the carryover effects. Two popular modeling approaches include Markov chain modeling and user modeling. In Markov chain modeling, treatment assignments in the past will affect the outcomes in the future through intermediate states that evolve as a Markov chain farias2022markovian. If the treatment assignments in the past do not directly affect the outcomes in the future other than going through the intermediate states, such intermediate states are referred to as “surrogates” athey2019surrogate, prentice1989surrogate. Similar to the overlap of network interference with marketplace interference, Markovian interference overlaps with switchback experiments given the temporal nature of conducting experiments. See glynn2020adaptive, hu2022switchback, jia2023faster. These aforementioned works focus on designing non-trivial experiments. Other works focus on conducting simple random experiments but combining them with non-trivial analysis. Some conduct experiments on one single unit liang2023randomization; some on multiple units athey2019surrogate, farias2022markovian, huang2023estimating, wen2024analysis, yang2020targeting. In user modeling, the treatment effects are modeled using specific agent level models. See hohnhold2015focusing, munro2023causal.
Violation of Assumption (ref) usually concerns violating the identical distribution assumption, which leads to many directions in the literature such as generalizability. This is an important topic in causal inference, yet we omit this topic from this manuscript. Instead, we discuss the literature on treatment heterogeneity. This literature usually uses covariates to better describe treatment heterogeneity and uses Assumption (ref) to overcome the violation of Assumption (ref). As we have seen in Section (ref), the value of covariates lies not only in reducing bias (through identifying all confounders) but also in reducing variance (through explaining variations coming from different covariates).
Using covariates in causal inference has a long history, probably as long as the causal inference literature itself. The earlier books of cochran1948experimental, cox2000theory have formally introduced classical designs of experiments, such as factorial design, block randomization, and stratified randomized design. As rubin2008comment commented, it is important to balance covariates in handling heterogeneity in randomized experiments. Two of the most popular ways to balance covariates are through stratified randomized design (Definition (ref) in Section (ref)) and through regression based methods (Section (ref)).
In line with stratified randomized design, there are numerous ways to partition units into strata. See cytrynbaum2021designing, greevy2004optimal, higgins2016improving, lu2011optimal, tabord2023stratification. One special case when time is modeled as a covariate, deng2013improving, jin2023toward, tang2020control, wu2022non study how to reduce variance in temporal experiments. These above works focus on designing non-trivial experiments. More generally, there are, of course, more extensive works that focus on conducting simple random experiments but combining them with non-trivial analysis such as matching, either directly on the covariates or on the propensity score or estimated propensity score. We are unable to survey this rich literature and only refer to a few papers such as abadie2012martingale, abadie2006large, bai2022optimality, bertsimas2015power, dehejia2002propensity, diamond2013genetic, hirano2003efficient, imai2009essential, kallus2018optimal, rosenbaum1983central, rosenbaum1984reducing, rosenbaum1989optimal, zubizarreta2012using, and references therein.
In line with regression based methods, Section (ref) is only a microcosm of the rich literature. We have only examined the $D_A$-optimal and $D$-optimal criteria in Section (ref). There are many more optimal criteria in the literature (e.g., A-optimal and E-optimal criteria), which take different perspectives to scalarize the covariance matrix (ref). We are unable to survey this rich literature and only refer to textbooks such as atkinson2007optimum, fedorov2013theory, pukelsheim2006optimal, silvey2013optimal, and papers such as alexanderian2014optimal, atkinson1975optimal, card1993minimum, de2019approximate, ruan2021linear, titterington1975optimal, ucinski2005t, and references therein. Solving the optimization problems under different optimal criteria is computationally challenging. One powerful tool to solve such optimization problems is semidefinite programming ahmadi2012convex, helmberg2002semidefinite, luo2010semidefinite, nie2014truncated, parrilo2003semidefinite, todd2001semidefinite, vandenberghe1996semidefinite. Recent works have also studied other objective functions that do not work with the covariance matrix (ref). See chattopadhyay2022balanced, harshaw2024balancing, li2015value, morris1979finite. The optimal experimental design literature can also be extended to stepped wedge designs where a treatment is sequentially rolled out over a number of time periods. See brown2006stepped, hemming2015stepped, hussey2007design, li2018optimal, xiong2023optimal.
Recently, as a result of the rise of clinical trial applications with sequential patient enrollment, covariate balancing problems are examined from the perspective of sequentially revealed covariate information. The adaptive clinical trial literature refers to this perspective as “covariate-adaptive” experiments, as the treatment assignments of future units depend on the covariates and treatment assignments of past units. The treatment assignments of future units do not depend on the observed outcomes of past units, so they are not “response-adaptive” hu2006theory. This literature stems from Efron's “biased coin” design efron1971forcing. Subsequently, pocock1975sequential and atkinson1982optimum develop general covariate-adaptive versions that make this literature popular. For recent developments, see atkinson1982optimum, atkinson1999optimum, bertsimas2019covariate, bhat2020near, kapelner2014matching, rosenberger2008handling, rosenberger2015randomization, zhao2024pigeonhole, and references therein.
Violation of Assumption (ref) leads to many directions in the literature. One direction among them is the literature on response-adaptive experiments, or adaptive experiments for short hu2006theory. In an adaptive experiment, the units are sequentially enrolled into the experiment. Following convention, unit $1$ arrives first, followed by unit $2$, and the last being unit $n$. In the setting with no covariate, $W_j$, the treatment assignment for each unit $j \in [n]$, depends on $Y_1, ..., Y_{j-1}$ the observed outcomes of the past units as well as $W_1, ..., W_{j-1}$, the past treatment assignments. Adaptive experiments are known to improve statistical efficiency and are thus desirable to experimenters murphy2005experimental, offer2021adaptive. However, adaptive experiments sometimes violate Assumption (ref). Consider the following two examples.
Examples (ref) and (ref) are only the two simplest examples. More generally, because $W_j$, the treatment assignment for a unit, is correlated with $Y_1,...,Y_{j-1}$, the observed outcomes of the past units, it is also correlated with $(Y_1(1), Y_1(0)), ..., (Y_{j-1}(1), Y_{j-1}(0))$, the potential outcomes of the past units (unless some additional assumptions are made), thus violating Assumption (ref). Because Assumption (ref) is violated, using the standard difference-in-means estimator can be biased. See bowden2017unbiased, deshpande2018accurate, dimakopoulou2021online, hadad2021confidence, hirano2023asymptotic, nie2018adaptively, shin2019sample, shin2019bias, zhan2021off, zhan2023policy, zhang2020inference, zhang2021statistical under various settings and various adaptive policies.
The phenomenon shown in Example (ref) is referred to as “optimistic sampling” bias in shin2019sample. To correct for such a bias in estimation, two typical solutions are concerned with using suitable estimators and designing suitable experiments. These two solutions can, of course, be combined. We next introduce Assumption (ref), the sequential random assignment assumption, which is usually satisfied under these two solutions.
\begin{assumption+}{(ref)$^{**}$}[Sequential random assignment] For each unit $j \in [n]$, conditional on the history $\big\{ Y_i, W_i \big\}_{i=1}^{j-1}$, the treatment assignment $W_j$ and the pair of potential outcomes $(Y_j(1), Y_j(0))$ are independent, that is,
\end{assumption+}
Starting from using suitable estimators, the IPW estimator (Definition (ref)) and the aggregate estimator (Definition (ref)) are two popular estimators to use. As long as Assumptions (ref) (i.e., the sequential random assignment assumption) and Assumption (ref) (i.e., the positivity assumption) hold, the IPW estimator is unbiased. See bojinov2019time, zhan2023policy, and references therein. The aggregate estimator requires that the experiment is conducted in stages, such as in Example (ref). See zhang2020inference and references therein. Additionally, the simulations literature has studied another type of estimation strategy when the potential outcomes are modeled to have specific parametric forms. See asmussen2007stochastic, glasserman2004monte, ross2013simulation for classical textbooks and references therein.
Moving on to designing suitable experiments, we start with Assumption (ref). Assumption (ref) is required for the IPW estimator to be well defined. As long as the IPW estimator is well defined, we can apply Assumption (ref) and martingale theorems to show that the IPW estimator is unbiased. One critical idea to make the IPW estimator unbiased is to enforce Assumption (ref). This can be done by enforcing an adaptive experiment to uniformly “explore” all the treatments with a strictly positive probability. Not only does this idea work for the IPW type of estimators, it also proves to be effective for other types of estimators as well. See bojinov2019time, bowden2017unbiased, dimakopoulou2021online, hadad2021confidence, ham2023designa, ham2023designb, zhan2021off, zhan2023policy, zhang2020inference, zhang2021statistical.
On the other hand, one could argue that many of the adaptive policies in the above works are not for the purpose of estimation and inference, but rather for a purpose called “(cumulative) regret minimization,” which is a partial cause of biases. We are unable to survey the rich literature on online learning, and only point to three alternative objectives called “best-arm identification” and “simple regret minimization” abbasi2018best, adusumilli2022minimax, audibert2010best, bubeck2009pure, chen2017adaptive, chen2018optimal, chen2023active, kasy2021adaptive, kato2022best, mannor2004sample, miao2023personalized, naby2024, qin2017improving, qin2022open, russo2016simple, tang2022offline, xu2023online, wu2022adaptive, zhang2024deep, zhou2014optimal, as well as “variance minimization” antos2010active, armstrong2022asymptotic, aznag2024active, blackwell2022batch, carpentier2011finite, carpentier2011upper, carpentier2012minimax, carpentier2015adaptive, dai2024clip, deep2023asymptotically, fontaine2021online, grover2009active, hahn2011adaptive, russac2021b, wei2023adaptive, wei2024fair, xiong2023optimal, zhao2023adaptive. All these objectives are highly related (under certain assumptions some of them are even equivalent), and are often referred to as the pure-exploration objective. When both the regret minimization objective and the pure-exploration objective are combined, there will be trade-offs athey2022contextual, bui2011committing, drugan2013designing, erraqabi2017trading, krishnamurthy2024proportional, qin2024optimizing, simchi2023multi, yao2021power, zhong2021achieving. See qin2024optimizing for an excellent survey of the literature and a state-of-the-art framework without covariates.
The phenomenon shown in Example (ref) is referred to as “stopping” bias in shin2019sample. To correct for such a bias in estimation, we could borrow the same ideas from the two solutions in correcting for the optimistic sampling bias. Motivated by improving statistical efficiency and reducing sample size, many works especially focus on studying how to stop the experiment as early as possible. These works are collectively referred to as the sequential testing literature. One classical technique that is popular in this literature is called the “law of iterated logarithm.” See siegmund2013sequential, wald2004sequential for classical textbooks, and bibaut2022near, cho2024peeking, jamieson2018bandit, johari2015always, johari2017peeking, lindon2022anytime, liang2023experimental, malenica2023anytime, ramdas2020admissible, ramdas2023game for recent developments under various settings.
This manuscript focuses on experimental design problems that arise in causal inference contexts as viewed through an optimization lens. We discuss three major frameworks of experimental design problems in details: the robust optimization framework, the stochastic optimization framework, and the deterministic optimization framework. The first framework models the uncertainty of potential outcomes to be more ambiguous, wherein the optimal design is usually random. Conversely, the second framework postulates distributional knowledge and the third framework postulates well-specified models to capture the uncertainty of potential outcomes, wherein the optimal design is usually deterministic.
These different frameworks reflect different ways to model the uncertainty governing the potential outcomes. For an experimenter to choose from one of the three frameworks, we recommend choosing the appropriate framework based on the uncertainty that the experimenter faces. We distinguish three cases. In the first case when we have strong knowledge to explain the uncertainty, it is appropriate to adopt the deterministic optimization framework. As an example given in rubin1978bayesian, this could be the case when an industrial experiment comparing manufacturing procedures may have strong knowledge about the relationship between covariates and potential outcomes. In the second case when we have some knowledge to describe a prior distribution of the potential outcomes, it is appropriate to adopt the stochastic optimization framework. This could be the case when we have collected historical data to estimate the prior distribution, and when we believe Assumption (ref) to hold; that is, there is no distributional shift. Finally, in the third case when we have little knowledge about the potential outcomes, it is appropriate to adopt the robust optimization framework. This could be the case when we have no historical data, or, as commented in wu1981robustness, because “the experimenter's knowledge about the (potential outcomes) model is never perfect.”
Depending on how we model uncertainty, it is conceptually simple to cast experimental design problems as optimization problems under one of the three frameworks. This manuscript presents, in consequence, a range of potential research opportunities to study experimental design problems for causal inference through an optimization lens.
\setcitestyle{numbers}