EconBase
← Back to paper

Finite Population Inference for Factorial Designs and Panel Experiments with Imperfect Compliance

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

57,096 characters · 14 sections · 30 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Finite Population Inference for Factorial Designs and Panel Experiments with Imperfect Compliance

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

\if01 {

center[center omitted — 35 chars of source]

} \fi

abstractThis paper develops a finite population framework for analyzing causal effects in settings with imperfect compliance where multiple treatments affect the outcome of interest. Two prominent examples are factorial designs and panel experiments with imperfect compliance. I define finite population causal effects that capture the relative effectiveness of alternative treatment sequences. I provide nonparametric estimators for a rich class of factorial and dynamic causal effects and derive their finite population distributions as the sample size increases. Monte Carlo simulations illustrate the desirable properties of the estimators. Finally, I use the estimator for causal effects in factorial designs to revisit a famous voter mobilization experiment that analyzes the effects of voting encouragement through phone calls on turnout.

{\it Keywords:} Causal Inference, Factorial Designs, Dynamic Causal Effects, Finite Population, Nonparametric.

\spacingset{1.8}

Introduction

In a seminal paper, imbensangrist showed that the instrumental variables (IV) estimand in cross-sectional settings with a binary treatment and a binary instrument can be interpreted as the Local Average Treatment Effect (LATE), defined as the average treatment effect for the subpopulation that has its treatment status shifted by an excluded instrument, the so-called compliers.

However, several applications of IV methods deviate from the canonical setting. For instance, it is common to see researchers using IV methods in settings which the outcome depends on multiple treatments. Prominent examples are factorial designs Gerber_Green_2000,nick and panel experiments stango with imperfect compliance. In such settings, it is well known that standard IV methods can lead to misleading conclusions regarding causal effects.

In this paper, I develop a novel nonparametric identification approach for causal effects in factorial designs and panel experiments. Using standard instrumental variables assumptions along with a exclusion restriction for treatments, I show that causal effects associated to sequences of treatments (a vector of factors of interest, or the path of treatments taken through time), can be identified by exploiting variations in the sequence of assignments associated to such treatments.

The intuition behind this result is simple: under the treatment exclusion restriction, if we are interested in the effect of a single treatment, we identify its causal effects for compliers of its assignment by exploiting variations in that assignment. If we are interested in the causal effects of a sequence of two treatments, we identify causal effects by exploiting variations in the four possible assignments in a difference-in-differences format. In the case of a sequence of three treatments causal effects are identified using triple-differences, and so on.

My approach takes a purely design-based perspective on uncertainty, which allows for potential outcomes models to remain unspecified. I propose nonparametric estimators for causal effects which are consistent over the randomization distribution and derive their finite population asymptotic distributions. I use the results to propose valid inference procedures, which are potentially conservative in the presence of heterogeneous causal effects.

I conduct Monte Carlo simulation studies to analyze the properties of these estimators. The results show that the estimators exhibit desirable performance in terms of bias and confidence interval coverage, both in the factorial design and the panel experiment settings. Finally, I revisit the work of nick and apply the estimators in a real life famous factorial experiment from the political science literature, which studies the effects of voting encouragements through phone calls in turnout.

This paper contributes to several strands of the causal inference literature. First, to the literature on factorial designs with imperfect compliance chengsmall,Blackwell2017,schochet,Blackwell2023. More specifically, I rely on the same assumptions as Blackwell2017 and Blackwell2023, but define a new, rich class of causal effects and provide a new nonparametric estimator.

Second, to the literature on dynamic causal effects in the tradition of ROBINS19861393,ROBINS1987139S and murphy, where potential outcomes in a given period depend on the path of treatments taken until that period. I focus on the case of imperfect compliance and the identification of lag-$p$ dynamic causal effects, defined by bojinovshephard. I extend the results of bojinov to the case of panel experiments with imperfect compliance. I motivate my approach for the identification of dynamic causal effects by showing that standard IV estimands in general do not have a straightforward causal interpretation in the presence of dynamic causal effects even in the case where the exclusion restriction for treatments hold. See Appendix C for a causal decomposition of the Wald estimand in a setting with two time periods.

Finally, the paper relates to the literature on design-based inference in IV settings (Imbens_Rubin_2015; kang; jonashesh, peterkyrill). While these papers focus solely on estimators for IV settings with a single treatment and a single instrument, I focus on settings with multiple treatments and instruments. While Blackwell 2023 also provides finite population inference procedures, this is the first paper that studies design-based inference in panel experiments with imperfect compliance, to the best of my knowledge.

The paper is organized as follows. In Section 2 I introduce the framework for factorial designs and provide the identification results, estimators and its asymptotic properties for this setting. In section 3 I proceed similarly, but focus on the case of panel experiments. In Section 4 I conduct Monte Carlo simulations to analyze the properties of the estimators introduced previously, and in Section 5 I apply the techniques developed in Section 2 to the voter mobilization empirical application. Section 6 concludes. The supplemental appendix contains the proofs of the theorems presented in this paper and the auxiliary lemmas.

Factorial Designs

Framework and Target Parameters

Consider a setting in which $N$ units are part of an experiment with $K$ binary factors with levels $\left\{0,1\right\}$, where 0 denote the untreated level and 1 denote the treated level. Let $Z_{i,k}$ be an indicator for units assigned to the treated level for factor $k$ and $D_{i,k}$ denote an indicator for units that uptake treatment in factor $k$. Index sets are compactly written as $[N]:=\left\{1,...,N\right\}$ and $[K]:=\left\{1,...,K\right\}$.

Imperfect compliance arises from the fact that not all units assigned to treatment in factor $k$ actually take treatment, and not all units assigned to control remain untreated, that is, for all $k\in[K]$, $D_{i,k}\neq Z_{i,k}$ for some $i\in[N]$.

For an individual $i$, I denote the vectors of treatment and assignments up to factor $k$ respectively as $D_{i,1:k}$ and $Z_{i,1:k}$. Vectors of treatments and assignments that do not contain elements for factors $k-p$ to $k$ are defined as $D_{i,-(k-p:k)}$ and $Z_{i,-(k-p:k)}$. For a given factor $k$, I denote the cross-section of treatments and assignments respectively as $D_{1:N,k}$ and $Z_{1:N,k}$. Collections over individuals and factors are denoted by $D_{1:N,1:K}$ and $Z_{1:N,1:K}$.

Without further restrictions, potential outcomes of unit $i$ are a function of the full sequence of treatments and assignments, $Y_{i}(d_{1:N,1:K},z_{1:N,1:K})$, and potential treatments are a function of the full sequence of assignments, $D_{i,k}(z_{1:N,1:K})$. The next assumptions are invoked for identification.

Assumption 1 (No-Spillovers Across Units): For all $i\in[N]$ and $k\in[K]$,

align*[align* omitted — 279 chars of source]

Assumption 1 imposes that the potential outcome of unit $i$ depends only on the treatments and assignments of unit $i$, ruling out the possibility of spillover of both treatments and assignments across units. Assumption 1 is usually referred to as the the Stable Unit Treatment Value Assumption rubin80.

One of the fundamental assumptions of IV settings is the exclusion restriction. Its version for factorial designs is stated below, alongside an additional exclusion restriction for the first stage:

Assumption 2 (Exclusion Restrictions): For all $i\in[N]$ and $k\in[K]$, (i) $Y_{i}(d_{1:K},z_{1:K})=Y_{i}(d_{1:K})$ and (ii) $D_{i,k}(z_{1:K})=D_{i,k}(z_{k})$.

The first part of Assumption 2 is the standard exclusion restriction for IV settings. It states that the sequence of assignments does not affect potential outcomes directly. Assignments only affect potential outcomes to the extent that they affect treatment choices.

The second part of Assumption 2 restricts how assignments affect treatment choices. states that treatment uptake on factor $k$ only depends on the treatment assignment for factor $k$, not other factors, and is widely invoked in factorial design settings with imperfect compliance Blackwell2017,Blackwell2023.

The fundamental behavioral assumption in IV settings is the monotonicity assumption, which is provided below.

Assumption 3 (Monotonicity): For all $i\in[N]$, $k\in[K]$, $D_{i,k}(1)\geq D_{i,k}(0)$.

Monotonicity states that for each factor, units assigned to treatment ($Z_{i,k} = 1$) are at least as likely to take that treatment as units assigned to control ($Z_{i,k} = 0$). Under Assumption 3, units can be divide into three groups at each period of time defined by how units treatment choice for factor $k$ relates to assignment $k$: Always-takers ($AT_{k}$), Never-takers ($NT_{k}$) and Compliers ($C_{k}$). Note that an individual can be part of a group in a given for a given factor but it does not need to be in the same group through the whole sequence of assignments. Assumptions 2 (ii) and 3 combined imply that there are $3^{K}$ compliance types in an experiment with $K$ factors.

Finally, it is assumed that assignments for all factors are completely randomized, in the sense that they are independent from potential outcomes and treatments, and from assignments to other factors. Let $\mathcal{F}_{i}:=\left\{Y_{i}(.), D_{i,1:K}(.)\right\}$. The assumption is formalized below:

Assumption 4 (Individualistic Assignments):For all $i\in[N]$ and $k\in[K]$,

equation*[equation* omitted — 126 chars of source]

Assumption 4 states that the assignment for factor $k$ for unit $i$ is independent from the assignment of other units, the assignment for other factors and potential quantities. It is readily satisfied in completely randomized experiments (where $\mathbb{P}\left ( Z_{i,k}|Z_{-i,k},Z_{i,-k},\mathcal{F}_{1:N} \right )=c_{k}\in \left(0,1\right)$ for all $i\in[N]$), but it also allows for more complex assignment mechanisms.

In the factorial design setting, causal effects are the differences in potential outcomes for unit $i$ associated to different sequences of treatments, which are defined as $\tau_{i}(d_{1:K},\tilde{d}_{1:K})=Y_{i}(d_{1:K})-Y_{i}(\tilde{d}_{1:K})$.

The number of potential outcomes grows exponentially with the number of factors. For the sake of tractability, I focus on the $k-p$ to $k$ joint causal effect:

equation*[equation* omitted — 154 chars of source]

which can be interpreted as the causal effect of taking a sequence of treatments from factor $k-p$ to $k$ versus an alternative sequence, keeping the sequence of all other factors fixed.

In settings with imperfect compliance, the average causal effects are only identifiable for the group of compliers. I define a target parameters for the group of compliers of the sequence. The local $k-p$ to $k$ causal effect

equation*[equation* omitted — 182 chars of source]

which is the $k-p$ to $k$ causal effect for individuals that are compliers for all treatments considered within the sequence of interest.

Identification

I show that joint causal effects can be identified after potential outcomes associated to different sequences of treatment are identified separately. Define the local $k-p$ to $k$ response function as

equation*[equation* omitted — 137 chars of source]

for any $\textbf{d}\in\left\{0,1\right\}^{p+1}$. I show that local response functions can be identified for all combinations of assignments. The local causal effect is then generated by taking the difference between two local response functions: $\tau_{C_{k-p:k}}(\textbf{d},\tilde{\textbf{d}})=m_{k-p:k}(\textbf{d})-m_{k-p:k}(\tilde{\textbf{d}})$.

Before introducing the estimand, I defined the adapted propensity scores from factors $k-p$ to $k$ as $\pi_{i,k-p:k}(\textbf{z})=\mathbb{P}\left(Z_{i,k-p:k}=\textbf{z}\right)$, which can be computed along the observed assignment sequences. The fundamental object used for identification is the Horvitz-Thompson estimand. For a generic random variable $R_{i,t}$, the Horvitz-Thompson associated to the assignment sequence $\textbf{z}\in\left\{0,1\right\}^{p+1}$ is

equation*[equation* omitted — 178 chars of source]

The Horvitz-Thompson-type quantities are combined in what I call a multiple-differences format. Below, I define an expression for the general multiple difference in means across assignment paths, which I will refer to hereafter as the $p+1$ Horvitz-Thompson multiple-difference ($\Delta^{p+1}_{HT}$).

Definition ($\Delta^{p+1}_{HT}$): Define the $p+1$-th difference of Horvitz-Thompson estimands across assignment sequences from period $k-p$ to period $k$ of a random variable $R$ as is such that for $p\geq 0$,

align*[align* omitted — 627 chars of source]

When $p=0$, the operator is simply a difference in means:

align*[align* omitted — 476 chars of source]

When $p=1$, the multiple differences operator takes the form of difference-in-differences:

{

align*[align* omitted — 1,270 chars of source]

}

When $p=2$ the operator takes the form of a triple-difference, and so on. Thus, the $p+1$ difference can be used to exploit all possible variations in assignment paths from $k$ to $k-p$, while keeping the rest of the assignments fixed.

In order to use the Horvitz-Thompson multiple-difference estimands, one must assume that there is common support for the adapted propensity score for all possible sequences of factors.

Assumption 5: There exists $C^{L}<C^{U}\in(0,1)$ such that or all $i\in[N]$, $k\in[K]$, $C^{L}<\pi_{i,k-p:k}(\textbf{z})<C^{U}$ for all $\textbf{z}\in\left \{ 0,1 \right \}^{p+1}$.

The theorem below shows that causal effects associated to a single factor can be identified using a Wald-type estimand and that local response functions associated to a treatment path from $k-p$ to $k$ can be identified exploiting variations in the path of assignments from $k-p$ to $k$ in the multiple-differences format.

theoremUnder Assumptions 1-5, \begin{equation*} \frac{\mathbb{E}\left [\frac{1}{N} \sum_{i=1}^{N}\frac{Z_{i,k}Y_{i}}{\pi_{i,k}(1)}-\frac{1}{N}\sum_{i=1}^{N}\frac{(1-Z_{i,k})Y_{i}}{\pi_{i,k}(0)}|\mathcal{F}_{1:N,-k} \right ]}{\mathbb{E}\left [\frac{1}{N} \sum_{i=1}^{N}\frac{Z_{i,k}D_{i,k}}{\pi_{i,k}(1)}-\frac{1}{N}\sum_{i=1}^{N}\frac{(1-Z_{i,k})D_{i,k}}{\pi_{i,k}(0)}|\mathcal{F}_{1:N,-k} \right ]}=\tau_{C_{k}}(1,0) \end{equation*} and \begin{equation*} \frac{\Delta^{p+1}_{HT}\left ( \mathbb{E}\left [ \frac{1}{N}\sum_{i=1}^{N}\frac{Y_{i}\mathbf{1}\left\{D_{i,k-p:k}=d \right\}\mathbf{1}\left\{Z_{i,k-p:k}=z \right\}}{\pi_{i,k-p:k}(z)}|\mathcal{F}_{1:N,-(k-p:k)} \right ] \right )}{\Delta^{p+1}_{HT}\left ( \mathbb{E}\left [ \frac{1}{N}\sum_{i=1}^{N}\frac{\mathbf{1}\left\{D_{i,k-p:k}=d \right\}\mathbf{1}\left\{Z_{i,k-p:k}=z \right\}}{\pi_{i,k-p:k}(z)}|\mathcal{F}_{1:N,-(k-p:k)} \right ] \right )}=m_{k-p:k}(\textbf{d}) \end{equation*} for any $\textbf{d}\in\left\{0,1\right\}^{p+1}$.

Theorem 1 shows that the causal effect of factor $k$ is identified by a standard two-stage Horvitz-Thompson estimand. In order to identify causal effects associated to multiple factors, local response functions are identified separately by exploiting variation in the sequence of assignments that affect the factors of interest in the multiple-differences format.

Estimation and Inference

In this section, I study the asymptotic properties of the estimator corresponding to the estimands introduced in Theorem 1.

I first derive the asymptotic properties of the factor $k$ causal effect estimator over the randomization distribution and then proceed with the asymptotic properties of the estimator for general $k-p$ to $k$ causal effects.

When it comes to the case where we are interested in the causal effect of a single treatment, potential outcomes do not need to be identified separately. I propose a simple two-stage Horvitz-Thompson (HT) estimator, which is built using the adapted propensity score.

The nonparametric estimator for $\tau_{i,k}(1,0)$ is

equation*[equation* omitted — 116 chars of source]

where

align*[align* omitted — 233 chars of source]

The local factor $k$ causal effect is then estimated by plugging in $\widehat{\tau}_{i,k}(1,0)$:

equation*[equation* omitted — 179 chars of source]

Theorem 2 shows that the estimator is consistent over the randomization distribution, and asymptotically normal as the population size grows larger.

theoremSuppose that potential outcomes are bounded and Assumptions 1-5 hold. Then, \begin{equation*} \frac{\sqrt{N}\left \{ \widehat{\overline{\tau}}_{C_{k}}(1,0)-\overline{\tau}_{C_{k}}(1,0) \right \}}{\sigma_{k}(1,0)}\overset{d}{\rightarrow} \mathcal{N}(0,1),\ as\ N\rightarrow\infty \end{equation*} where $\sigma_{k}(1,0)$ is defined in the Appendix.

The variances of $\widehat{\overline{\tau}}_{C_{k}}(1,0)$ is the appropriate averages of the variance of $\widehat{\overline{\tau}}_{i,k}(1,0)$, which is generally not estimable as it depends on individual potential outcomes and potential treatments under both treatment and counterfactual. However, Lemma 1 in Appendix B shows that the variance of the reduced form and the first stage are bounded from above by a term that is estimable.

For hypothesis testing, I propose the estimation of a conservative bloom confidence interval, built with estimates of the upper bound of the variance of the reduced form, and the square of the estimate of the first stage.

The upper bound for the variance of the estimator of the reduced form $\widehat{\tau}_{i,k}^{RF}(1,0)$ is

equation*[equation* omitted — 180 chars of source]

and it can be consistently estimated by $\left ( \widehat{\gamma}_{i,k}^{RF}(1,0) \right )^{2}=\frac{Y_{i}^{2}\left\{ Z_{i,k}+(1-Z_{ik})\right\}}{\pi_{i,k}(Z_{i,k})^{2}}$ and plugged-in for the estimate of the confidence interval. The resulting $1-\alpha$ confidence intervals for the local factor $k$ causal effect is

equation*[equation* omitted — 260 chars of source]

bloom intervals exhibit good performance in terms of coverage rates for compliance rates greater than 10%. See kang for a thorough discussion about inference using IV in cross-sectional settings with a single binary factor.

For the general $k-p$ to $k$ local causal effect, I propose the separate estimation of local response functions through a multiple-differences Horvitz-Thompson type of estimator. That is, $\widehat{\tau}_{i,k-p:k}(\textbf{d},\tilde{\textbf{d}})=\widehat{m}_{i,k-p:k}(\textbf{d})-\widehat{m}_{i,k-p:k}(\tilde{\textbf{d}})$, where $\widehat{m}_{i,k-p:k}(\textbf{d})=\frac{\widehat{m}_{i,k-p:k}^{RF}(\textbf{d})}{\widehat{m}_{i,k-p:k}^{FS}(\textbf{d})}$ with

align*[align* omitted — 440 chars of source]

Plugging the unit $i$ estimates for the $k-p$ to $k$ local response function leads to the estimate of the average local response function, which is

equation*[equation* omitted — 188 chars of source]

Estimates for local $k-p$ to $k$ causal effects are generated by the difference of the estimates for the response functions associated to alternative treatment paths. Theorem 3 shows that the estimators are consistent over the randomization distribution, and asymptotically normal as the population size grows larger.

theoremSuppose that potential outcomes are bounded and Assumptions 1-5 hold. Then, \begin{equation*} \frac{\sqrt{N}\left \{ \widehat{\overline{\tau}}_{C_{k-p:k}}(d,\tilde{d})-\overline{\tau}_{C_{k-p:k}}(d,\tilde{d}) \right \}}{\sigma_{k-p:k}(d,\tilde{d})}\overset{d}{\rightarrow} \mathcal{N}(0,1),\ as\ N\rightarrow\infty \end{equation*} where $\sigma_{k-p:k}(\textbf{d},\tilde{\textbf{d}})$ is defined in the Appendix.

For hypothesis testing, conservative bloom confidence intervals are constructed using the upper bound for the variance of the reduced forms for the response functions and estimates of the first stage for the response functions.

The upper bound for the variance of the estimator of the reduced form is

align*[align* omitted — 557 chars of source]

and it can be consistently estimated and subsequently plugged-in to construct the confidence interval.

In Section 4.1, I conduct Monte Carlo simulations to study the asymptotic properties of these estimators.

Panel Experiments

Framework and Assumptions

Consider a balanced panel in which $N$ units are observed over $T$ periods of time. For each unit $i\in[N]$ and time period $t\in[T]$, we observe a binary instrumental variable $Z_{i,t}\in\left \{ 0,1 \right \}$, a binary treatment status $D_{i,t}\in\left \{ 0,1 \right \}$ and a real-valued scalar outcome $Y_{i,t}$. Let $\mathcal{F}_{1:N,1:t}$ denote the filtration generated by $Z_{1:N,1:t}$ and the panel of potential quantities. Note that since we are taking a design-based approach, conditioning on $Z_{i,1:t}$ is the same as conditioning on $Z_{i,1:t}$ and the potential quantities associated to such assignment path.

Without further restrictions, potential outcomes of unit $i$ in period $t$ are a function of the full panel of treatments and assignments, $Y_{i,t}(d_{1:N,1:T},z_{1:N,1:T})$, and potential treatments are a function of the full panel of assignments, $D_{i,t}(z_{1:N,1:T})$. The next assumptions are invoked for identification.

Assumption 6 (No-Spillovers and No-Anticipation): For all $i\in[N]$, $t\in[N]$,

align*[align* omitted — 300 chars of source]

Assumption 6 imposes that the potential outcome of unit $i$ in period $t$ depends only on the treatment and assignment paths of unit $i$ until period $t$, ruling out the possibility of spillover of both treatments and assignments across units, as well as future treatments affecting past potential outcomes. It also imposes that the potential treatment from unit $i$ in period $t$ depends only on the assignment paths of unit $i$ until period $t$. To put it shortly, Assumption 6 imposes both SUTVA and no-anticipation.

The dynamic version of the exclusion restriction is stated below, alongside an additional exclusion restriction for the first stage.

Assumption 7 (Exclusion Restrictions): For all $i\in[N]$ and $t\in[T]$, (i) $Y_{i,t}(d_{1:t},z_{1:t})=Y_{i,t}(d_{1:t})$ and (ii) $D_{i,t}(z_{1:t})=D_{i,t}(z_{t})$.

The first part of Assumption 7 is the standard exclusion restriction for dynamic IV settings. It states that the path of assignments does not affect potential outcomes directly. Assignments only affect potential outcomes to the extent that they affect treatment choices.

The second part of Assumption 7 adapts the treatment exclusion restriction from factorial designs to the panel experiment setting. It imposes that potential treatments in period $t$ depend only on the instruments in period $t$, formalizing the intuition that assignment is “targeted” towards a single treatment.

Assumption 8 (Monotonicity): For all $i\in[N]$ and $t\in[T]$, $D_{i,t}(1)\geq D_{i,t}(0)$.

Monotonicity states that at each period units assigned to treatment ($Z_{i,t}=1$) are at least as likely to take treatment as units assigned to control ($Z_{i,t}=0$). Under Assumption 3, units can be divide into three groups at each period of time defined by how units treatment choice in period t relates to treatment assignment in period t: Always-takers ($AT_{t}$), Never-takes ($NT_{t}$) and Compliers ($C_{t}$). Note that an individual can be part of a group in a given period, it does not need to be in the same group through the whole path of assignments.

Assumption 9 is the standard exogeneity assumption for assignments in panel experiments

Assumption 9 (Individualistic and sequentially randomized assignment): For all $i\in[N]$, $t\in[T]$ and $z_{1:N,1:t-1}\in \left\{ 0,1\right\}^{N\times (t-1)}$,

equation*[equation* omitted — 204 chars of source]

Assumption 9 is the exogeneity assumption from bojinovshephard and bojinov, and imposes that conditional on its own past assignments, treatments and outcomes, the assignment for unit $i$ at time $t$ is independent of the past assignments and outcomes of all other units as well as all other contemporaneous assignments.

Define the adapted propensity score for the panel experiment setting as

equation*[equation* omitted — 191 chars of source]

The common support assumption under which it is a valid building block for the estimands is stated below:

Assumption 10 (Common Support): There exists $C^{L}<C^{U}\in(0,1)$ such that or all $i\in[N]$, $t\in[T]$, $C^{L}<\pi_{i,t-p:t}(\textbf{z})<C^{U}$ for all $\textbf{z}\in\left \{ 0,1 \right \}^{p+1}$.

Define the dynamic causal effect of a treatment path versus an alternative treatment path in period $t$ as $\tau_{i,t}(d_{i,1:t},\tilde{d}_{i,1:t})=Y_{i,t}(d_{i,1:t})-Y_{i,t}(\tilde{d}_{i,1:t})$.

The number of potential outcomes grows exponentially with the periods of time. For the sake of tractability, it is common to focus on lag-p dynamic causal effect as defined in bojinov. For $0\leq p\leq t$, and $\textbf{d},\tilde{\textbf{d}}\in\left \{ 0,1 \right \}^{p+1}$, the lag-p dynamic causal effect is defined as

equation*[equation* omitted — 154 chars of source]

which can be interpreted as the causal effect of taking a treatment path from period $t-p$ to $p$ versus an alternative path, keeping the path until period $t-p-1$ fixed.

The local time $t$ lag-$p$ dynamic causal effects and the time $t$ local lag-$p$ response functions are defined, respectively, as

align*[align* omitted — 302 chars of source]

Next, I show that the quantities are also identified using a multiple-differences Wald type of estimand.

Identification

Similar to the case of factorial designs, in panel experiments with imperfect compliance where exclusion restrictions for the treatments hold, potential outcomes associated to sequences of treatments are identified by exploiting variations in the corresponding sequence of assignments using the multiple-differences operator for Horvitz-Thompson quantities. Theorem 4 shows that the local lag-0 dynamic causal effect is identified under a two-stage Horvitz-Thompson estimand, and that local lag-$p$ response functions are identified using the multiple-differences approach.

theoremUnder Assumptions 6-10, \begin{equation*} \frac{\mathbb{E}\left [\frac{1}{N} \sum_{i=1}^{N}\frac{Z_{i,t}Y_{i,t}}{\pi_{i,t}(1)}-\frac{1}{N}\sum_{i=1}^{N}\frac{(1-Z_{i,t})Y_{i,t}}{\pi_{i,t}(0)}|\mathcal{F}_{1:N,t-1} \right ]}{\mathbb{E}\left [\frac{1}{N} \sum_{i=1}^{N}\frac{Z_{i,t}D_{i,t}}{\pi_{i,t}(1)}-\frac{1}{N}\sum_{i=1}^{N}\frac{(1-Z_{i,t})D_{i,t}}{\pi_{i,t}(0)}|\mathcal{F}_{1:N,t-1} \right ]}=\tau_{C_{t},t}(1,0;0) \end{equation*} and \begin{equation*} \frac{\Delta^{p+1}_{HT}\left ( \mathbb{E}\left [ \frac{1}{N}\sum_{i=1}^{N}\frac{Y_{i,t}\mathbf{1}\left\{D_{i,t-p:t}=d \right\}\mathbf{1}\left\{Z_{i,t-p:t}=z \right\}}{\pi_{i,t-p:t}(z)}|\mathcal{F}_{1:N,t-p-1} \right ] \right )}{\Delta^{p+1}_{HT}\left ( \mathbb{E}\left [ \frac{1}{N}\sum_{i=1}^{N}\frac{\mathbf{1}\left\{D_{i,t-p:t}=d \right\}\mathbf{1}\left\{Z_{i,t-p:t}=z \right\}}{\pi_{i,t-p:t}(z)}|\mathcal{F}_{1:N,t-p-1} \right ] \right )}=m_{t-p:t}(\textbf{d}) \end{equation*} for any $\textbf{d}\in\left\{0,1\right\}^{p+1}$.

Theorem 4 shows that the mechanics for identifying dynamic causal effects and local lag-$p$ response functions is the same as the one for identifying causal effects and local response functions in factorial designs. That is because we assume the treatment exclusion restriction holds in both settings. Ultimately, the multiple-difference approach works because compliance types assignment-specific (factor-specific in the factorial design and period-specific in the panel experiment design), and thus shifts in an assignment identifies potential quantities for individuals who comply to that assignment while keeping the response fixed for all other treatments. Therefore, the multiple-differences approach allows for potential quantities for compliers of a sequence of assignments to be sequentially identified by sequentially exploring variations in such sequence of assignments while keeping other assignments fixed.

Estimation and Inference

The estimator procedure for local dynamic causal effects and local dynamic response functions is similar to the one proposed in Section 2.3. The nonparametric estimator for $\tau_{i,t}(1,0;0)$ is

equation*[equation* omitted — 122 chars of source]

where

align*[align* omitted — 238 chars of source]

The time-t lag-0 dynamic causal effect can be estimated by plugging in $\widehat{\tau}_{i,t}(1,0;0)$:

equation*[equation* omitted — 187 chars of source]

Theorem 5 shows that the estimator is consistent over the randomization distribution, and asymptotically normal as the population size grows larger.

theoremSuppose that potential outcomes are bounded. Under Assumptions 6-10, \begin{equation*} \frac{\sqrt{N}\left \{ \widehat{\overline{\tau}}_{C_{t},t}(1,0;0)-\overline{\tau}_{C_{t},t}(1,0;0) \right \}}{\sigma_{t}(1,0;0)}\overset{d}{\rightarrow} \mathcal{N}(0,1),\ as\ N\rightarrow\infty \end{equation*} where $\sigma_{t}(1,0;0)$ is defined in Section 5 of Appendix A.

The variance of $\widehat{\overline{\tau}}_{C_{t},t}(1,0;0)$ is the appropriate averages of the variance of $\widehat{\overline{\tau}}_{i,t}(1,0;0)$, which are generally not estimable as they depends on individual potential outcomes and potential treatments under both treatment and counterfactual. However, Lemma 4 in Appendix B shows that the variance of the reduced form and the first stage are bounded from above by a term that is estimable.

For the general lag-$p$ dynamic causal effect, I propose the separate estimation of lag-$p$ dynamic response functions through a multiple-differences Horvitz-Thompson type of estimator. That is, $\widehat{\tau}_{i,t}(\textbf{d},\tilde{\textbf{d}};p)=\widehat{m}_{i,t}(\textbf{d})-\widehat{m}_{i,t}(\tilde{\textbf{d}})$, where $\widehat{m}_{i,t}(\textbf{d})=\frac{\widehat{m}_{i,t}^{RF}(\textbf{d})}{\widehat{m}_{i,t}^{FS}(\textbf{d})}$ with

align*[align* omitted — 449 chars of source]

Plugging the unit $i$, period $t$ estimates for the lag-$p$ local response function as leads to estimates of the time-$t$ local lag-$p$ response function:

equation*[equation* omitted — 176 chars of source]

The appropriate lag-p dynamic causal effects are generated by the difference of the estimates for the response functions. Theorem 6 shows that the estimator is consistent over the randomization distribution, and asymptotically normal as the population size grows larger.

theoremSuppose that potential outcomes are bounded. Under Assumptions 6-10, \begin{equation*} \frac{\sqrt{N}\left \{ \widehat{\overline{\tau}}_{C_{t-p:t},t}(d,\tilde{d};p)-\overline{\tau}_{C_{t-p:},t}(d,\tilde{d};p) \right \}}{\sigma_{t}(d,\tilde{d};p)}\overset{d}{\rightarrow} \mathcal{N}(0,1),\ as\ N\rightarrow\infty \end{equation*} where $\sigma_{t}(\textbf{d},\tilde{\textbf{d}};p)$ is defined in Section 6 of Appendix A.

Once again, I use bloom confidence intervals for hypothesis testing. In Section 4.2, I conduct Monte Carlo simulations to study the asymptotic properties of these estimators.

Monte Carlo Simulations

In this section I show the desirable finite-sample properties of the proposed nonparametric estimators for the factors $k-p$ to $k$ causal effects and the local lag-$p$ dynamic causal effects. I study their finite-sample properties in terms of the average bias (Av. Bias), median bias (Med. Bias), root mean-squared error (RMSE), coverage of the Confidence Interval (Cover) and the Confidence Interval length (CIL).

Factorial Designs

I consider a setting with $N=1.000$ units and $K=2$ binary factors. For each factor, assignment is randomized following a Bernoulli distribution:

equation*[equation* omitted — 123 chars of source]

I set $p_{i,1}=p_{1}=0.5$ and $p_{i,2}=p_{2}=0.5$, such that the data generating process emulates a completely randomized experiment. Potential outcomes are generated according to the following linear model:

equation*[equation* omitted — 117 chars of source]

I set $\beta_{0}=0, \beta_{1}=0.5,\beta_{2}=1,\beta_{1,2}=0.25$ and $\varepsilon_{i}\sim \mathcal{N}(0,1)$. Potential treatments $D_{i,k}(z_{k})$ are generated following Bernoulli distributions with parameters $\delta_{k}(z_{k})$. For $k=1$ I set $\delta_1(z_{1})=0.2+0.7z_{1}$, such that compliance rate is 70%. For $k=1$ I set $\delta_2(z_{2})=0.2+0.6z_{2}$, such that compliance rate is 60%. I then conduct 10.000 Monte Carlo experiments to evaluate the performance of the Horvitz-Thompson estimators proposed in Section 2.3, the results are summarized in Table 1.

table[table omitted — 896 chars of source]

The first three columns present the results for the estimators of factors 1 to 2 causal effects associated to three different treatment uptakes using the multiple-differences approach for response functions separately and then generating the estimate of the causal effect by taking the difference across response functions. Note that in the case of $K=2$, the multiple-differences estimand takes the form of a two-stage difference-in-differences across assignments.

The estimator exhibits negligible bias for the three causal effects and Bloom confidence interval have empirical coverage close to the desired 95% coverage, with lengths fairly stable across the different effects.

The last two columns present the results for the two-stage Horvitz-Thompson estimator of the factor $k$ specific causal effects. Again, the estimator exhibits negligible bias across the simulation and Bloom confidence interval have a good performance in terms of coverage. Confidence interval have approximately have the length of the confidence intervals from the multiple-differences estimator.

Panel Experiments

I consider a balanced panel setting with $N=1000$ and $T=2$. The simulation focuses on the lag-p dynamic causal effects with $p=t$, that is, the dynamic causal effects associated to the whole treatment path in the setting. Assignment is sequentially randomized following a Bernoulli distribution:

equation*[equation* omitted — 181 chars of source]

I set $p_{i,t}=p_{i}=0.6$. For the choice model. Outcomes in period $t$ are specified to have the following linear working model:

equation*[equation* omitted — 92 chars of source]

I set $\beta_{t}=1$ and $\beta_{t-1}=0.5$, the vector $\beta_{1:t-2}$ to be a vector of zeros, and $U_{i,t}(D_{1:t})\sim \mathcal{N}(0,1)$. Potential treatments at period $t$ are generated in a way that the compliance rate is 50% at each time period.

In Table 2, I compare the performance of the proposed Horvitz-Thompson estimator (HT) with the performance of the period-specific 2SLS estimator\footnote{In Section C of the Appendix I present a causal decomposition of the period-specific 2SLS estimand and show that it does not hold a straightforward causal interpretation in $t=2$.}, taking the local lag-0 dynamic causal effect as the target parameter.

table[table omitted — 1,120 chars of source]

The first two columns show the results for the first period. When $t=1$, dynamics play no role in the model. Hence, both the nonparametric estimator and the static 2SLS show little to none Monte Carlo bias. Moreover, the coverage is close to the desired 95%, with the 2SLS estimator showing a tighter Confidence Interval on average. When it comes to the second period, the Horvitz-Thompson estimator remains consistent in the Monte Carlo exercise, as shown by the third column. The period-specific 2SLS estimator is severely biased, with coverage far from the desired 95%. The last two columns stack the lag-0 estimates across the two time periods. Thus, the results can be interpreted as a weighted average of the time-$t$ results, which explains why the performance of the 2SLS estimator is better than in the second period alone. As the number of periods grows larger, however, one should expect the 2SLS estimator to perform increasingly worse with the number of time periods.

In Table 4, I present the simulation results for different time-$t$ local lag-1 dynamic causal effects. The lag-1 effects can only be estimated for period 2. I present results for the difference-in-differences modified Horvitz-Thompson estimator (HT) and a multivariate 2SLS estimator (MV2SLS). I consider the performance of the estimators with respect to effects of full exposure in the first two columns, exposure in the second period in the third and fourth columns, and exposure in the first period in the last two columns.

The multivariate 2SLS specification yields substantially biased estimates for the three different target parameters. Coverage is closer to the desired 95% than in the simulations for the period-specific 2SLS. However, it is never greater than 81.5%.

The proposed nonparametric estimator exhibits great Monte Carlo performance. The average bias, median bias and root mean-squared error for the causal effects are small, and the coverage of the conservative confidence interval is close to the desired 95% coverage. Confidence interval lengths are fairly stable across the considered treatment paths.

table[table omitted — 1,066 chars of source]

Overall, the Monte Carlo Simulations assert the desirable finite-sample performance of the proposed estimators for dynamic causal effects over the randomization, while bringing evidence of pitfalls associated to the standard 2SLS methods in the presence of time-varying heterogeneity.

Application - GOTV Experiment

In this section, I use the estimators introduced in Section 2 to revisit the causal effects of get-out-the-vote (GOTV) efforts on youth voting turnout in the large field experiment conducted through the Youth Vote Coalition by nick.

In this experiment individual registered voters were randomly assigned to receive a call from a volunteer phone bank ($Z_{i,1}$), a professional phone bank ($Z_{i,2}$), both or neither. The probabilities of assignment for $Z_{i,1}$ and $Z_{i,2}$ are 0.5 each, which means the four possible paths of assignment are observed with the same probability. The outcome of interest is turnout in the 2004 presidential election. Imperfect compliance in this setting arises from the fact that not all individuals actually picked up the phone when received the call.

In order to estimate the causal effects of the phone call encouragements on turnout, I follow Blackwell2017 and focus on college-age respondents from sites in which the experiment was implemented successfully (both forms of contact were assigned and there are no violations of the exclusion restriction), leading to $N=26,974$ respondents. Table 4 summarizes the results.

table[table omitted — 751 chars of source]

The first three rows of Table 4 show that volunteer and professional phone bank calls had a positive impact on turnout. The causal effect of receiving both calls vs not receiving any call is similar in magnitude to the causal effect of receiving only one type of calls vs no calls, which suggests that these two interventions are not complementary. The last two rows present the factor $k$ specific causal effects, which are far smaller in magnitude. The fact that the estimates for $\tau_{C_{1}}(1,0)$ and $\tau_{C_{2}}(1,0)$ are smaller than the estimates for $\tau_{C_{1:2}}(\left\{1,0\right\},\left\{0,0\right\})$ and $\tau_{C_{1:2}}(\left\{0,1\right\},\left\{0,0\right\})$ suggest that there are negative effect of the interaction of treatments, indicating diminishing returns to GOTV efforts, which has been previously documented by Blackwell2017.

With the exception of $\tau_{C_{2}}(1,0)$, all the estimates for causal effects are statistically significant. I use Bloom confidence intervals for hypothesis testing. The joint compliance rates are estimated to be around 22.5%, which indicates that Bloom CIs should exhibit coverage rate close to the desired 95%. The estimates for the joint compliance rate are stable across the modified first stages, as displayed in Table 5. When $K=2$, the modified first-stage has expectation $\pm\frac{1}{N}\left | C_{1:2}\right |$, where the negative signs in the two middle rows arise mechanically from the multiple-differences operator, as shown in Section 1 of Appendix A. Thus, the estimates for compliance rates should be interpreted as the absolute values displayed in Table 5.

table[table omitted — 634 chars of source]

Overall, the result of the application align to the previous results regarding the 2004 Youth Vote Coalition GOTV experiment and show evidence of diminishing returns to follow-up contact on turnout. The fundamental difference between the previous results and the results presented in this Section comes from the approach towards uncertainty. While nick and Blackwell2017 construct confidence intervals based on a traditional sampling perspective, I think of the population of the experiment as fixed, and the phone call assignments as stochastic. The framework from Section 2 implies that if Assumptions 1-5 hold for this finite population, then the CIs from Table 4 be interpreted as a valid, but possibly conservative 95% confidence intervals for the causal effects of interest.

Conclusion

This paper develops a design-based framework for experimental settings with imperfect compliance in which multiple treatments can affect the outcome of interest. I focus on the cases of factorial designs and panel experiments and show that under standard instrumental variable assumptions and a treatment exclusion restriction which can be justified in both settings, causal effects associated to sequences of treatments can be identified by exploiting variation in the sequence of assignments associated to it in what I define as a two-stage multiple-differences estimand.

I introduce nonparametric estimators for causal effects which are consistent over the randomization distribution, and derive their finite population asymptotic distributions. I illustrate the desirable properties of the proposed estimators through Monte Carlo simulation studies and apply the estimator for factorial designs in a famous political science setting. The results show that the causal effects of GOTV efforts are large and statistically significant under potentially conservative inference procedures.