Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,221 characters · 17 sections · 30 citation commands
Causal Inference in Counterbalanced Within-Subjects Designs
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Carryover effects, Sequential exchangeability, Potential outcomes framework, Treatment effect identification, Within-subject experimental design, Causal inference assumptions
\spacingset{1.45}
Experimental designs are increasingly used in the social and behavioral sciences in estimating the causal effects of various interventions. The randomized controlled trial (RCT) is widely regarded to be the `gold standard' when it comes to estimating the causal effects of treatment as opposed to estimates from an observational study. Typically, participants are divided into distinct treatment and control groups, allowing for straightforward comparisons of intervention outcomes in RCTs.\footnote{This design is called between-subjects design in some fields since interventions are randomly assigned between subjects 02_maxwell2017designing.} However, this conventional approach is not the only experimental design used.
Within-subjects design, also known as repeated-measures design, offers an alternative that exposes all participants to every condition in a study. This design is sometimes used in fields such as experimental psychology keren2014between, management sciences Lane2024, political science 09_Clifford, and even medical trials Sarkies2019, which enables researchers to compare responses across treatments for the same individuals. Experimenters often employ these methods when they face logistical challenges in recruitment, implementation (particularly regarding the costs associated with specific trials), or when the population concerned is relatively small, and the experimenter wishes to ensure there is common support in terms of the exposure of the treatment with respect to the covariates (02_maxwell2017designing; ErlebacherAlbert1977Daao). However, this experimental design introduces challenges in causal inference.
For example, imagine a study testing the effect of background music on concentration. Participants complete two problem-solving tasks: one in silence and one with music. If every participant always does the silent condition first and the music condition second, any observed difference in performance might not be due to the music itself but rather to other factors. For instance, they might perform better in the second task simply because they have become more familiar with the problem format (a practice effect) or worse due to mental fatigue. Since these order-related influences are confounded with the treatment, the researcher cannot tell whether any observed changes are truly caused by the music or just by the passage of time and repeated testing. The chief culprit confounding the estimate of the treatment effect in within-subjects designs are carryover effects, which are the residual effects from the first period that influence responses in the second period 02_maxwell2017designing.
To address these concerns, counterbalancing is often applied to within-subjects designs. In clinical trials, this design is also known as crossover trials or AB/BA designs ZHANG2022_crossovertrials, sibbald1998understanding, matthew_multiperiod_crossover. This approach randomizes the order of treatment and control conditions for participants, assuming that this would cancel out any carryover effects, resulting in an unbiased estimate. Figure (ref) shows how a counterbalanced within-subjects design is implemented in the case where there are two experimental conditions. However, this assumption is often overly optimistic. Counterbalancing assumes that carryover effects are symmetric and cancel out, yet this symmetry is unverifiable a priori. Differential carryover effects, where one condition's impact persists or interacts with subsequent conditions in an asymmetric manner, can bias estimates of the average treatment effect even when participants are properly counterbalanced.
There are some advantages in using within-subjects design in some empirical settings, as it aligns more naturally with economic and psychological theories that model how individuals respond to changing conditions. For example, in empirical studies testing utility theories and preference, participants might exposed to multiple conditions of the study, allowing the experimenter to observe how an individual makes such a tradeoff 09_Charness.
Despite their inferential limitations, counterbalanced within-subjects designs offer several practical advantages. They allow researchers to collect multiple observations per participant, improving statistical efficiency—especially useful when working with small populations or when recruitment is costly or difficult, such as in studies involving venture capitalists Lane2024. In such cases, a between-subjects design may suffer from issues like limited overlap or lack of common support across covariates, making within-subjects designs an attractive alternative despite their potential identification problems.
In this paper, we analyze the implications of using counterbalanced within-subjects designs for causal inference, analyzing the problem with the potential outcomes framework. Specifically, we first explore the assumptions required in counterbalanced within-subjects design to produce treatment effect estimates consistent with those obtained from between-subjects design in the case of two discrete time periods, with the goal of estimating the Average Treatment Effect (ATE) as defined in the potential outcomes framework. We then introduce sequential exchangeability as a generalized assumption required, alongside the standard assumptions of the Rubin Causal Model, to identify the causal effects of treatment in other sequential designs similar to a counterbalanced within-subjects design. We introduce sequential randomization and selective sequential randomization in Section (ref) as alternative sequential designs where the assumptions of sequential exchangeability is more credible than the commonly used counterbalanced within-subjects design.
Sequential designs are often used to improve the efficiency and reliability of experiments. For example, 04_whitehead2020estimation use a design that allows dropping underperforming treatments mid-trial, 08_zhou2018sequential propose a rerandomization method to improve covariate balance over time, and 03_tamura2011estimation apply a sequential parallel design to reduce placebo effects. While these approaches differ in motivation, they all rely on using data collected across stages to better estimate treatment effects. We focus on how similar ideas apply to counterbalanced within-subjects designs and relate sequential exchangeability as a condition under which such designs can identify the average treatment effect (ATE).
We want to emphasize that such assumptions are typically difficult to verify a priori and hence, we do not advocate the use of such experimental designs to estimate ATE, even in the case of sequential randomization where the assumptions are more credible. However, in situations where such a design is required due to the aforementioned logistical challenges, we provide a way to reason, and guidance to ensure that the treatment effects are properly estimated.
Our work builds on existing work of analyzing experimental designs within the potential‐outcomes framework such as in mediation analysis, factorial designs, and conjoint studies, by making explicit the precise causal assumptions each design invokes. In mediation analysis, for example, researchers decompose a treatment’s total effect into direct and indirect pathways under a sequential ignorability assumption that links potential mediators and outcomes imai2013experimental,11_imai2010general. Conjoint analysis, meanwhile, treats each attribute as a component “treatment,” isolating marginal contributions and interactions via randomized component assignment and carefully articulated exclusion restrictions egami2019causal. In the same spirit, our paper examines counterbalanced within-subjects design to ask: what is the set of causal assumptions required to point‐identify the ATE in counterbalanced within‐subject designs? By iterating these identifying assumptions in full generality, we extend the potential‐outcomes framework to reason for arbitrary, higher‐order carryover and learning effects, rather than relying on the narrow parametric or symmetry conditions typical of two‐period crossover estimators 23_AnalysisOfCrossoverTrial.
The paper is organized as follows. Section (ref) outlines the identification assumptions required in a counterbalanced within-subjects design and provides illustrative examples of when these assumptions might be violated. Section (ref) generalizes the assumptions and introduce the concept of sequential exchangeability, as well as alternative sequential designs where the assumptions are more credible. Section (ref) discusses heuristic checks to verify whether violations of the assumptions occur and explores ways to recover an unbiased causal estimate using simulations under certain additional assumptions. Section (ref) concludes the paper.
In this section, we formalize the identification challenges that arises from counterbalanced within-subjects design by comparing it with an experiment using a between-subjects design, with the goal of recovering the Average Treatment Effect (ATE). We develop the additional assumptions required so that the estimates in a counterbalanced within-subjects design will yield the same results as in a between-subjects design in this section.
Assume that we have $n$ units of observations, where we have the set of associated observed outcome $\{Y_i\}_{i=0}^n$ and an treatment $\{Z_i\}_{i=0}^n$. In the potential outcomes framework, we imagine jointly observing both outcomes of treatment and control 10_holland1986statistics. We denote the potential outcomes as $(Y_i(1), Y_i(0))$, where $1$ and $0$ indicate that the unit received treatment or control, respectively. From these definitions, we can define the key identity
which captures the relationship between potential outcomes and observed outcomes. Note that by this identity, $Y_i$ denotes the observed outcome of unit $i$. We also observe the pre-treatment covariates, denoted as $\{X_i\}_{i=0}^n$.
Per the standard definition of the Average Treatment Effect (ATE), it is defined as the difference in the potential outcomes of treatment and control\[\tau = \mathbb{E}[Y_i(1) - Y_i(0)]\] Under the standard assumptions of the Rubin Causal Model, the Average Treatment Effect is identified if 10_holland1986statistics
These assumptions together imply that the unobserved counterfactuals can be recovered, in expectation, from observed outcomes conditional on covariates. This leads to the identifiability of the Average Treatment Effect in a between-subjects design where the assumptions are plausibly satisfied.
One of the biggest appeals of an experimental design such as a between-subjects design is that participants are randomly assigned to receive either treatment or control but not both. This removes selection bias if done properly, allowing us to reasonably assume that ignorability holds, and therefore, a simple difference in means estimator without adjustment would lead to an unbiased estimate of ATE.
We begin the analysis of a counterbalanced within-subjects design in the case where there are two experimental conditions, treatment and control, and two discrete time periods, $t_1$ and $t_2$. In a counterbalanced within-subjects design, each unit of observation is assigned to both treatment and control, but with each unit being randomly assigned to the sequence of treatment. Let $S_i$ be the sequence a unit $i$ is assigned to, where $\{S_i\}^n_{i=1}$ for when $S_i = 0$, unit $i$ receives control at $t_1$ and then receives treatment at $t_2$ and when $S_i=1$ unit $i$ receives treatment at $t_1$ and then receives control at $t_2$.
$Z_{i,t_1} \in \{0,1\}$ indicates that unit $i$ receives control or treatment at $t_1$, where $Z_{i,t_1} = 1$ when unit $i$ receives treatment and 0 otherwise. In a counterbalanced within-subjects design, $Z_{i,t_2} = 1 - Z_{i,t_1}$ and $S_i = Z_{i,t_1}$. We can define the observed outcomes $Y_{i, t_1}^{\text{obs}}$ and $Y_{i, t_2}^{\text{obs}}$ in terms of the potential outcomes as
At $t_1$, the design is equivalent to a between-subjects design, with the potential outcomes at $t_1$, denoted by $Y_{i,t_1}(1)$ for treatment and $Y_{i,t_1}(0)$ for control. The average treatment effect (ATE) at $t_1$ is defined as \[ \tau_{t_1} = \mathbb{E}\big[Y_{i,t_1}(1) - Y_{i,t_1}(0)\big]. \] Since participants in a counterbalanced within-subjects design are randomized into the sequence in which they receive the treatment, $\tau_{t_1} = \tau$ trivially as a result.
At $t_2$, counterbalanced within-subjects design deviates from the standard framework. As a result of the sequential treatment design, participants switch conditions so that the observed potential outcomes are:
Importantly, the potential outcomes $Y_{i,t_2}(0,0)$ and $Y_{i,t_2}(1,1)$ are never observed for all $i$ in a counterbalanced within-subjects design. Thus, the “average treatment effect” at $t_2$, denoted as $\tau_{t_2}$, in a counterbalanced within-subjects design is defined as \[ \tau_{t_2} = \mathbb{E}\big[ Y_{i,t_2}(0,1) - Y_{i,t_2}(1,0) \big]. \] Note that this is not the typical parameter that people usually think of when they say ATE, since this is more accurate to describe it as the average treatment effect at $t_2$ given the treated was untreated at time $t_1$ and the untreated was treated at time $t_1$.
Unlike $\tau_{t_1}$, the equality $\tau_{t_2} = \tau$ does not hold automatically. In the following section, we develop the assumptions required such that \[ \tau_{t_2} = \mathbb{E}\big[ Y_{i,t_2}(0,1) - Y_{i,t_2}(1,0) \big] = \mathbb{E}\big[ Y_{i}(1) - Y_{i}(0) \big] = \tau \] since this is often the purpose of using counterbalanced within-subjects design in causal inference. In the following sections, we will primarily focus on the analysis of $\tau_{t_2}$.
To justify our focus on \(\tau_{t_2}\), consider a common estimation method used to estimate ATE in a counterbalanced within‐subjects design that employs the following regression specification with an OLS estimator:
A key motivation for using a counterbalanced within-subjects design is to increase effective sample sizes by pooling observations across time periods. In this context, the pooled OLS estimate \(\hat{\alpha}_1\) is interpreted as the estimated average treatment effect in a counterbalanced within-subjects design, that is, \(\hat{\tau}^{\text{CWSD}} = \hat{\alpha}_1\).
Proposition (ref) shows that the pooled OLS estimate from a counterbalanced within-subjects design is a weighted average of the treatment effects from each period. This decomposition allows us to interpret the estimator as a blend of the two time-specific effects, with the weight \(q\) depending on the variance structure of the treatment assignment across periods.
This comes immediately from the linearity of expectations, which justifies our subsequent focus on the analysis of $\tau_{t_2}$.
One of the major challenges in estimating the average treatment effect (ATE) in a counterbalanced within-subjects design is the presence of carryover effects. These effects refer to residual influences from the treatment (or control) administered in the first period, denoted by \(z_{t_1}\), on the outcome observed at the second period, beyond the effect of the treatment received at \(z_{t_2}\). Such residual effects can introduce unwanted variability or bias in the estimation of the treatment effect at time \(t_2\).
Figure (ref) illustrates the causal pathways in a counterbalanced within-subjects design. The principal identification challenge arises at \(t_2\), where the experimental design allows treatment received at \(t_1\) to confound the effect of the treatment at \(t_2\).
The first potential source of bias is due to the direct carryover effects of the treatment at \(t_1\) on the outcome at \(t_2\). This effect is captured by the causal pathway \[ Z_{i,t_1} \to Y_{i,t_2}, \] indicating that any residual impact of the treatment (or control) at \(t_1\) can directly influence the observed outcome \(Y_{i,t_2}^{\text{obs}}\) at \(t_2\). Before we formalize the concept of direct carryover effects, we first have to define the potential outcome $Y_{i,t_2}(z_{i,t_2})$.
Note that in a counterbalanced within-subjects design, this is never observed. However, it serves as a useful counterfactual for reasoning about what the outcome would have been if unit $i$ appeared at $t_2$ and was only assigned to $z_{i,t_2}$ without any prior exposure. An implicit assumption in the way $Y_{i,t_2}(z_{i,t_2})$ is defined is that \[Y_{i,t_2}(z_{i,t_2}) \hspace{-.2em}\perp\!\!\!\!\perp\hspace{-.2em} Z_{i,t_1} \mid Z_{i,t_2}, X_{i,t_2}\] since $Z_{i,t_1}$ `does not exist' in this counterfactual scenario. With this, we can define direct carryover effects as follows:
In words, the direct carryover effects as stated in Definition (ref) measures the difference between the potential outcome when an individual receives the sequence \((z_{t_1}, z_{t_2})\) and the potential outcome if only the treatment at \(t_2\) were administered. In the perspective of the potential outcomes framework, we can imagine that unit $i$ was recruited at time $t_2$ without participating at all at time $t_1$, and has the potential outcome $Y_{i,t_2}\bigl(z_{t_2}\bigr)$.
In a counterbalanced within-subjects design, counterbalancing is typically implemented under the assumption that these direct carryover effects are simple—that is, they are symmetric and cancel out in expectation 02_maxwell2017designing. Formally, we define simple direct carryover effects as follows.
In addition to direct effects, indirect carryover effects may occur when the treatment at \(t_1\) influences covariates measured at \(t_2\), which in turn affect the outcome \(Y_{i,t_2}\). As shown in Figure (ref), there are two potential indirect pathways:
To ensure an unbiased estimate of the treatment effect at \(t_2\), these indirect pathways must be blocked, typically by controlling for \(X_{i,t_2}\) or with randomization at $t_2$.
Finally, given that outcomes may change over time irrespective of treatment, it is necessary to account for time-specific shifts. In other words, even in the absence of treatment, the evolution of outcomes between \(t_1\) and \(t_2\) should be similar across treatment and control groups. This is analogous to the common parallel trends assumption in difference-in-differences analysis, and is stated formally as \[ Y_{i,t_2}(z_{i,t_2}) = Y_{i,t_1}(z_{i,t_2}) + c, \quad \forall z_{i,t_2} \in \{0,1\}, \quad c \in \mathbb{R}. \] This assumption implies that any time-related change in the outcome is the same for both treatment and control conditions, differing only by a constant \(c\).
The direct and indirect carryover effects, along with the parallel trends assumption, comprise the additional assumptions required for the identification of the ATE in a counterbalanced within-subjects design. We summarize these conditions in the following theorem.
If we assume that direct carryover effects follow the “simple" definition in Definition (ref), then we have \(\mathbb{E}[C_{i,0 \to 1}] = \mathbb{E}[C_{i,1 \to 0}]\). This balance allows us to obtain an unbiased estimate of the treatment effect at time \( t_2 \), as shown in Theorem (ref).
However, when carryover effects are not simple — meaning \(\mathbb{E}[C_{i,0 \to 1}] \neq \mathbb{E}[C_{i,1 \to 0}]\) — counterbalancing alone cannot solve the issue. To illustrate why, consider the following (somewhat extreme) example.
Perder Pharma, a pharmaceutical company, is developing a pain-relief drug called Moxycontin, which is opioid-based and highly effective at reducing pain (denoted as some \(\tau < 0\)). To save costs, they use a counterbalanced within-subjects design to test its effect on back pain, randomizing participants into different sequences: some receive Moxycontin first, while others get it later.
For those who receive a placebo first, the carryover effect \(C_{0 \to 1}\) is close to zero—these participants are fine until they later take Moxycontin, which genuinely helps with pain. However, the carryover effect in the other direction, \(C_{1 \to 0}\), is much larger than \(|\tau|\) and greater than zero. This is because Moxycontin is an opioid, and withdrawal causes severe pain, making things much worse for participants who are taken off the drug. The withdrawal is so extreme that the FDA ultimately concludes Moxycontin increases back pain and denies Perder's application. As a result, more than 400,000 lives are saved, and many others avoid the devastation of opioid addiction.
Beyond direct carryover effects, counterbalanced within-subjects design also lacks full randomization in the second time period, unlike between-subjects design. This is because the treatment condition in the first period, \(Z_{i,t_1}\), can indirectly influence covariates—whether observed or unobserved—at \(t_2\). If these covariates (\(X_{i,t_2}\)) are not accounted for, the assumption of ignorability at \(t_2\) may be violated. To illustrate how this assumption may fail to hold, we provide two examples (1) induced selection bias and (2) survivorship bias.
In the case of an induced selection bias, consider an experiment testing the effect of Ozempic on blood sugar levels. A key covariate influencing blood sugar is physical activity. If participants receiving Ozempic are motivated to increase their physical activity---perhaps due to perceived health improvements or weight loss---while the control group’s activity levels remain unchanged, this introduces imbalance. At time \(t_2\), the Ozempic group may, on average, have higher physical activity levels, which would confound the estimated treatment effect and violate the assumption of ignorability.
As for survivorship bias, consider an experiment testing a treatment for cancer patients. At \(t_1\), participants in the control group who have worse health outcomes might die, effectively dropping out at \(t_2\), whereas some in the treatment group who would have died if assigned to control at \(t_1\) might survive and be observed at \(t_2\). This leads to participants in the treatment group being relatively healthier at \(t_2\) compared to their counterparts in the control group, assuming that the participants in the treatment group at $t_2$ does not suffer a decline of health due to not receiving the treatment at $t_1$.
Lastly, the parallel trends assumption may also be violated. Suppose a study involves participants solving complex problems in two sessions, using a counterbalanced design. If receiving an innovative problem-solving prompt in the first session not only improves performance in that session but also fundamentally changes how participants approach future problems, their learning trajectory shifts. As a result, the outcome evolution between sessions is nonparallel, violating the parallel trends assumption.
In this section, we develop the concept of sequential exchangeability, extending the original assumptions stated in Theorem (ref). This extension is motivated by the observation that carryover effects are not unique to counterbalanced within-subjects design but can also influence inference in other sequential experimental designs.
For instance, in 09_Clifford, the authors advocated for the adoption of pre-post designs, wherein participants are initially exposed to a control condition before transitioning to a between-subjects design. They posited that this approach enhances precision. However, in one of the experiments they replicated, carryover effects appeared to influence inference—not necessarily in the direction of the effect, but rather in its magnitude.
Experiment 2 in 09_Clifford replicated a landmark study on foreign aid by GilensMartin2001PIaC, investigating the influence of factual information regarding federal budget allocation for foreign aid on public opinion. In that study, respondents were informed that spending on foreign aid constituted less than 1% of the federal budget, after which they were asked whether foreign aid spending should be increased or decreased, with responses recorded on a five-point scale.
To compare different experimental designs, participants were randomly assigned to one of two conditions: the post-only design or the pre-post design. In the post-only design, participants were randomized to receive the budgetary information, consistent with a between-subjects approach. Conversely, in the pre-post design, participants first provided their opinion on foreign aid without the budgetary information, were subsequently exposed to the budgetary information, and finally responded to the same question again. The results of this replication indicate substantial differences in the estimated effects, as depicted in Figure (ref).
This example demonstrates how experimental design choices can markedly affect inference. Specifically, it suggests that the estimated treatment effect varies with the design, potentially due to carryover effects. Formally, we observe that: \[ \tau^{\text{between-subjects}} = \mathbb{E}[Y_i(1) - Y_i(0)] \neq \tau^{\text{pre-post}} = \mathbb{E}[Y_i(0,1) - Y_i(0,0)] \] Thus, the notion of sequential exchangeability provides a framework to examine the conditions under which these two estimands may be equivalent.
In the counterbalanced within-subjects design design, the sequence in which participants receive the treatment or control is randomized such that $Z_{i,t_1}$ determines $Z_{i,t_2}$; specifically, $Z_{i,t_1} = 1 - Z_{i,t_2}$. In contrast, the pre-post design employed in Experiment 2 in 09_Clifford randomizes participants' treatment status at time $t_2$, with all participants receiving the control in time $t_1$.
Moreover, other sequential experimental designs randomize treatment assignment in both time periods, as will be discussed further in Section (ref). In order to generalize, we can redefine Equation (ref) as follows:
We now formally define sequential exchangeability as follows:
The generalization to $n$ time periods and $n$ conditions can be found in the Appendix.
The primary distinction between sequential exchangeability and the assumptions required to identify the average treatment effect (ATE) in counterbalanced within-subjects design lies in Assumption (ref). Intuitively, if carryover effects can be appropriately controlled—such that the potential outcome at $t_2$ corresponds to that which would be observed if the participant were recruited at $t_2$ without any prior intervention. We want to note however that this is a strong assumption that can be easily violated. We explore some of the practical challenges in controlling for the direct carryover effects in Section (ref).
Evaluating Experiment 2 in 09_Clifford with respect to Assumption (ref), it is plausible that Assumptions (ref) and (ref) hold, given that participants were randomized at $t_2$ and the parallel trend assumption seem plausible. However, due to the nature of the experiment, administering the question at $t_1$ may introduce carryover effects that are challenging to control for. The mechanism by which treatment at $t_1$ influences the outcome at $t_2$ is ambiguous, thereby potentially violating Assumption (ref).
Based on the sequential exchangeability assumptions, we propose several alternative experimental designs to counterbalanced within-subjects design that may be more appropriate in contexts where the underlying assumptions are more plausible. Although these designs have been previously employed for diverse purposes as noted in the introduction, our focus here is on their utility for estimating the average treatment effect (ATE).
A notable advantage of the sequential randomization design depicted in Figure (ref) is its ability to incorporate \(Z_{i,t_1}\) as a control variable without introducing perfect collinearity. In traditional counterbalanced within-subjects design settings, the deterministic relationship \(Z_{i,t_1} = 1 - Z_{i,t_2}\) renders it infeasible to include both treatment indicators in an OLS regression model. By contrast, the sequential randomization design circumvents this issue, thereby allowing more robust estimation strategies. Nonetheless, a key challenge remains in accurately characterizing the nature of carryover effects; the risk of misspecification looms large if these effects are not thoroughly understood.
The most significant benefit of the proposed experimental designs is the credibility of the ignorability assumption at time \(t_2\) (Assumption (ref)). When randomization is implemented independently in the second stage, the potential for unobserved confounding—particularly through indirect carryover effects—is substantially reduced. This is especially valuable since experimental designs are often applied to estimate the average treatment effect, without assuming how potential covariates might confound the estimate.
Sequential experimental designs can and should be carefully tailored to the specific context of an experiment. For instance, consider the selective sequential randomization design illustrated in Figure (ref). If carryover effects are negligible in one of the conditions—such as when the control involves a placebo (e.g., a sugar pill assumed to have no effects of the outcome of interest)—then it may be reasonable to relax Assumption (ref), where the assumption is true for \(z_{i,t_1} = 0\) but not \(z_{i,t_1} = 1\). In this scenario, the selective design could yield reliable ATE estimates. However, if the control represents a current standard treatment and the experimental arm involves a novel intervention, where both conditions may have significant carryover effects, relaxing the assumption would be inappropriate, as the differential impact of prior exposure could bias the estimation of ATE. This emphasis on accounting for the nature of the experiment and treatment, and on adapting the randomization of treatment assignment for causal identification, echoes the rationale behind dynamic treatment regimes and Sequential Multiple Assignment Randomized Trial (SMART) designs, which similarly account for sequential assignments to estimate causal effects under minimal modeling assumptions 25_SMARTDesign.
For completeness, we state the following corollary linking sequential exchangeability to the validity of the proposed experimental designs.
The proof proceeds analogously to that of Theorem (ref).
Carryover effects from the treatment at time $t_1$ can be viewed as a form of omitted variable bias—essentially, a `covariate' that needs to be properly accounted for. As with any covariate, misspecifying its functional form can lead to biased estimates. If carryover effects are linearly separable, then a fixed-effects model that includes treatment status at $t_1$ as a control is typically sufficient. However, when carryover effects interact with the subsequent treatment at $t_2$, estimation becomes more complex.
Consider two illustrative models where carryover effects cannot be fully addressed by simple fixed-effects approaches. If the carryover effects from $t_1$ interacts outcomes at $t_2$ in an additive way, we can model the outcome as:
Here, the effect of the second treatment $Z_{i,t_2}$ depends additively on whether the first treatment $Z_{i,t_1}$ was received, through the interaction term $\gamma (Z_{i,t_1} \cdot Z_{i,t_2})$.
If, instead, the prior treatment amplifies or dampens the effect of the later treatment in a nonlinear way—as in the following model:
In this formulation, the strength of the effect of \( Z_{i,t_2} \) is modulated by prior exposure \( Z_{i,t_1} \) through an exponentiated function. When \( \gamma > 0 \), prior treatment increases the effective coefficient on \( Z_{i,t_2} \), amplifying its influence on the outcome. Conversely, when \( \gamma < 0 \), the coefficient is dampened.
Figure (ref) shows the bias of the estimated treatment effect, $\beta_1$, at $t_2$ varying the magnitude of the carryover effect, $\gamma$, as specified in Model 1 and Model 2. These two examples show how misspecification or ignoring the correct functional form of carryover can bias estimates. In Model 1, adding a simple interaction term can often account for such dependence, whereas in Model 2 requires a different approach to properly estimate the true effect of the second treatment.
23_AnalysisOfCrossoverTrial provides an analysis of estimating counterbalanced within-subjects design (referred to as 2×2 crossover trials) in the case of no carryover effects, as well as additive and interactive carryover effects, using OLS and REML estimators. However, as noted, the functional form of the carryover effect can be difficult to discern, and misspecification can lead to biased estimates.
Although counterbalanced within‐subjects designs present some identification challenges, there are practical strategies available to researchers to address these problems. This section outlines five approaches—Fisher's Exact Test, heuristic checks, the use of a washout period, covariate adjustment, and using Sequential Randomization as described in Section (ref)—that can help detect violations of the assumptions or satisfy them needed for consistent estimation of the causal effect.
Fisher's Exact Test is a nonparametric method widely used to assess the independence between two categorical variables, making it particularly useful in small sample settings. In a counterbalanced within-subjects design, if there is concern that assumption violations may be contaminating the analysis, one strategy is to discard data from $t_2$ and focus exclusively on $t_1$. By doing so, researchers can use Fisher's Exact Test at $t_1$ to determine whether the treatment assignment is independent of the observed outcomes.
However, this approach does not address the issue of limited overlap and lack of common support, which may arise with the sample sizes typically used in counterbalanced within-subjects design. Therefore, while it is a viable option, the experimenter may still need to ensure covariate balance at $t_1$, and its applicability may be limited in certain contexts.
A straightforward diagnostic is to compare the estimated treatment effect in the first period, \(\hat{\tau}_{t_1}\), to the estimated effect in the second period, \(\hat{\tau}_{t_2}\). A marked discrepancy between these estimates may indicate violations of the no‐carryover or no‐time‐drift assumptions (analogous to pre‐trend checks in a difference‐in‐differences setting).
When large differences emerge, they serve as a clear warning sign that carryover or other time‐based confounders may be at play. However, consistency between \(\hat{\tau}_{t_1}\) and \(\hat{\tau}_{t_2}\) does not guarantee the validity of the requisite assumptions: random sampling variability could mask carryover effects, while strongly offsetting biases might also produce superficially similar estimates. Inconsistencies between the estimates also do not imply that the assumptions are violated as differences might arise due to sampling variability. As such, this check is neither necessary (i.e., estimates may differ for benign reasons) nor sufficient (i.e., matching estimates could still reflect compensating biases).
Despite these limitations, heuristic checks remain a useful first step in diagnosing possible threats to identification and can guide further analyses. However, given that counterbalanced within-subjects design is often employed due to a small sample size, in practice, it can be difficult to determine that difference in estimates is driven by the violation of the assumptions, or sampling variation, or both.
A more targeted design‐based remedy is to incorporate a washout period between treatments. This approach is particularly relevant when the treatment is expected to have direct physiological or psychological effects that persist over time, but decay with a sufficient gap between $t_1$ and $t_2$. By lengthening the gap between \(t_1\) and \(t_2\), the influence of the initial treatment on the second‐period outcome may subside, thereby reducing or eliminating any carryover effects.
Researchers should draw on subject‐matter knowledge (e.g., pharmacokinetics, psychological adaptation timelines) to determine the length of time needed for treatment effects to dissipate. Even with an adequate washout period, however, confounding due to the effect of treatment on the covariates remains a concern. If, for instance, the initial treatment induces behavioral or physiological changes that persist, mere passage of time may not fully remove indirect carryover pathways. Thus, the potential for lingering confounders should still be carefully addressed (e.g., by measuring and adjusting for relevant characteristics at \(t_2\)).
If the carryover effect is entirely mediated by observable covariates at the second time period, \(X_{i,t_2}\) (i.e., if all carryover effects operate indirectly through these covariates), then adjusting for \(X_{i,t_2}\) fully accounts for any such indirect influence. Formally, controlling for the covariates at $t_2$ satisfies Assumption (ref).
To illustrate with the Ozempic experiment, suppose that any direct carryover effects of the treatment on blood sugar levels at \(t_1\) become negligible following an appropriate washout period; then the assumption of no direct carryover effects becomes more credible. However, if participants assigned to Ozempic continue to maintain elevated physical activity relative to those in the control group, a counterbalanced within-subjects design will violate Assumption (ref).
To account for this in a counterbalanced within-subjects design, we can adapt methods from observational causal inference by adjusting for $X_{i,t_2}$ using regression models or propensity score weighting. These approaches block indirect pathways through which prior treatment might influence the outcome. In contrast, a simple time dummy only captures average period effects and fails to account for treatment–covariate interactions. Nevertheless, while such adjustments are theoretically viable in a counterbalanced within-subjects design, they may be difficult to implement in practice, and model misspecification can introduce bias. Therefore, even when direct carryover effects seem unlikely or negligible, we recommend using sequential randomization as described in Section (ref).
To demonstrate this, we simulated a counterbalanced within-subjects design where participants were randomized to sequences, and varied the effect of the effect of $Z_{t_1}$ on $X_{t_2}$ from –10 to 10. We estimated the treatment effect using four methods: a correctly specified model, a propensity score method, a fixed time effect model, and an unadjusted model.
As shown in Figure (ref), only the correctly specified and propensity score models consistently recovered the true treatment effect. The fixed effect and unadjusted models produced biased estimates, with bias increasing alongside the treatment’s effect on the covariates. We repeated the simulation using sequential randomization, which yielded unbiased estimates regardless of model (mis)specification.
Although adjusting for covariates at $t_2$ can help block indirect carryover pathways, care must be taken not to condition on previous treatment outcomes such as \( Y_{i,t_1} \). In particular, consider the path \[ Z_{i,t_2} \leftarrow Z_{i,t_1} \rightarrow Y_{i,t_1} \leftarrow X_{i,t_1} \rightarrow X_{i,t_2} \rightarrow Y_{i,t_2} \] $Y_{i, t_1}$ is a collider and conditioning on it opens a non-causal backdoor path from \( Z_{i,t_1} \) to \( Y_{i,t_2} \). This violates the backdoor criterion and introduces what is known as collider bias or M-bias 21_M_Bias, 22_good_and_bad_crtl. As a result, estimates of the causal effect may become biased. Thus, in a counterbalanced within-subjects design, treatment outcomes at $t_1$ should not be conditioned on when estimating causal effects of \( Z_{i,t_2} \) on $Y_{i,t_2}$.
Our analysis of counterbalanced within-subjects designs shows the potential challenges in using this methodology for causal inference. While this design can offer advantages—such as increased statistical power, more precise individual-level comparisons, and efficient data collection—it also introduces complications that can bias treatment effect estimates if not carefully accounted for. The primary concern stems from carryover effects, which may persist asymmetrically across treatment conditions, and violations of stability over treatment and time, which can confound treatment estimates across sequential periods.
Through a formal treatment of these issues within the potential outcomes framework, we demonstrate that counterbalancing does not inherently eliminate biases introduced by differential carryover effects. Our findings suggest that, although counterbalancing is often assumed to mitigate such issues, it relies on the strong and unverifiable assumption that carryover effects are symmetric and cancel out. When this assumption does not hold treatment effect estimates may be substantially biased.
Furthermore, issues that arise in counterbalanced within-subjects design also appear in other sequential experimental designs where the estimand of interest is the ATE. We formalize the notion of sequential exchangeability, which applies to alternative experimental designs to counterbalanced within-subjects design where the assumptions are more plausibly met. However, we caution that despite the improvements, the difficulty of controlling for carryover effects can substantially hinder the estimation of the ATE and is not a perfect substitute for between-subjects design.
To address these concerns, we propose several practical strategies that researchers can employ when implementing counterbalanced within-subjects designs. First, heuristic checks comparing treatment effects across different time periods can serve as an initial diagnostic tool to detect potential violations of key assumptions. Second, incorporating a washout period between conditions can reduce the impact of lingering treatment effects, particularly in studies involving physiological or psychological interventions. Finally, adjusting for covariate changes between periods can help correct for biases arising from time-dependent confounders that may be influenced by prior treatment exposure, assuming that all carryover effects can be explained indirectly via the covariates.
Despite its limitations, counterbalanced within-subjects design can be a valuable tool for experimental research participant recruitment is constrained. However, researchers should assess whether counterbalanced within-subjects design is the most appropriate design choice for their specific study context, or whether alternative methodologies—such as between-subjects designs or sequential randomization—might provide more reliable estimates of causal effects.