Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,364 characters · 14 sections · 76 citation commands
Inference for Matched Tuples and Fully Blocked Factorial Designs
KEYWORDS: Randomized controlled trials, matched tuples, matched pairs, multiple treatments, factorial designs
JEL classification codes: C12, C14 \hypersetup{pageanchor=false} \thispagestyle{empty} \hypersetup{pageanchor=true} \setcounter{page}{1}
This paper studies inference in randomized controlled trials with multiple treatments, where treatment status is determined according to a “matched tuples” design. If there are $|\mathcal{D}|$ possible treatments, then by a matched tuples design, we mean an experimental design where units are sampled i.i.d.\ from the population of interest, grouped into “homogeneous” blocks of size $|\mathcal{D}|$, and finally, within each block, exactly one individual is randomly assigned to each of the $|\mathcal{D}|$ treatments. As such, matched tuples designs generalize the concept of matched pairs designs to settings with more than two treatments. Matched tuples designs are commonly used in the social sciences: see bold2018experimental, brown2020inducing, McKenzie2013, and McKenzie2014 for examples in economics, and are often motivated using the simulation evidence presented in bruhn2009pursuit. However, we are not aware of any formal results which establish valid asymptotically exact methods of inference for matched tuples designs. Accordingly, in this paper we establish general results about estimation and inference for matched tuples designs, and then apply these results to study the asymptotic properties of what we call “fully-blocked” $2^K$ factorial designs.
We first study estimation and inference for matched tuples designs in the general setting where the parameter of interest is a vector of linear contrasts over the collection of average outcomes for each treatment. Parameters of this form include standard average treatment effects (ATEs) used to compare one treatment relative to another, but as we explain below also include more complicated parameters which may be of interest, for instance, in the analysis of factorial designs. We first establish conditions under which a sample analogue estimator is asymptotically normal and construct a consistent estimator of its corresponding asymptotic variance. Combining these results establishes the asymptotic validity of tests based on these estimators. We then consider the asymptotic properties of two commonly recommended inference procedures. The first is based on a linear regression with block fixed effects. Importantly, we find the $t$-test based on such a regression is in general not valid for testing the null hypothesis that a pairwise ATE is equal to a prespecified value. The second is based on a linear regression with cluster-robust standard errors, where clusters are defined at the block level. Here we find that the corresponding $t$-test is generally valid but conservative, and that this conservativeness increases in the number of treatments.
Next, we apply our results to study the asymptotic properties of “fully-blocked” $2^K$ factorial designs. Factorial designs are classical experimental designs wu2011experiments which are increasingly being used in the social sciences alatas2012targeting, besedevs2012age, dellavigna2016voting, kaur2015self, karlan2014agricultural. In a $2^K$ factorial design, each treatment is a combination of multiple “factors,” where each factor can take two distinct values, or “levels.” As a consequence, a full $2^K$ factorial design can be thought of as a randomized experiment with $2^K$ distinct treatments (importantly however, the analysis of factorial designs typically considers factorial effects as the parameters of interest: see Section (ref) for a definition). A fully-blocked factorial design is then simply a matched tuples design with blocks of size $2^K$. Leveraging our previous results, we establish that our estimator achieves a lower asymptotic variance under the fully-blocked design than under any stratified factorial design which stratifies the experimental sample into a finite number of “large” strata (such designs include complete randomization as a special case). We also consider settings where only one factor may be of primary interest, and establish that even in such cases it is more efficient to perform a fully-blocked design than to perform a matched pairs design which exclusively focuses on the primary factor of interest.
In a simulation study, we find that although our inference results are asymptotically exact, our proposed tests may be conservative in finite samples when the experiment features many treatments or many blocking variables. Accordingly, we also study the behavior of a matched tuples design with “replicates,” where we form blocks of size two times the number of treatments, and each treatment is assigned exactly twice at random within each block. Although we find that such a design results in an estimator with slightly larger mean-squared error, the rejection probabilities of our proposed tests become much closer to the nominal level, which may result in improved power. Further discussion is provided in Section (ref) below.
Although the analysis of matched tuples designs has to our knowledge not received much attention, there are large literatures on both the analysis of matched pairs designs and the analysis of factorial designs. Recent papers which have analyzed the properties of matched pairs designs include athey2017econometrics, bai2021inference, bai2022optimality, chaisemartin2022at, cytrynbaum2021designing, imai2009essential, jiang2020bootstrap, fogarty2018regression, and van2012adaptive. Our analysis builds directly on the framework developed in bai2021inference, and our Theorems (ref) and (ref) nest some of their results when specialized to the setting of a binary treatment. cytrynbaum2021designing considers a generalization of matched pairs designs, a special case of which he refers to as a matched tuples design. However, his design groups units into homogeneous blocks in order to assign a binary treatment with unequal treatment fractions. In contrast, we consider a setting where units are grouped into homogeneous blocks in order to assign multiple treatments.
Recent papers which have analyzed factorial designs include rubin2016, dasgupta2015causal, li2020rerandomization, muralidharan2019factorial, pashley2019causal, and liu2022randomization. Our setup and notation for $2^K$ factorial designs mirrors the framework introduced in dasgupta2015causal, although our setup differs in that we consider a “super-population” framework where potential outcomes are modeled as random, whereas they maintain a finite population framework where potential outcomes are modeled as fixed.\footnote{The finite population “design-based" perspective may be particularly attractive in settings where the experimental sample is not explicitly drawn from a larger population. In Appendix (ref) we provide some preliminary simulation evidence that our proposed estimators may be relevant in such a setting as well, however, given the simulation evidence in chaisemartin2022at and our currently incomplete understanding of the design-based properties of our estimators, we do not make any general claims in this paper.} Borrowing the framework from dasgupta2015causal, rubin2016 and li2020rerandomization propose re-randomization designs for factorial experiments which are shown to have favorable efficiency properties relative to a completely randomized design. Although we do not provide formal results comparing our fully-blocked design to these re-randomization designs, our simulation evidence suggests that, at least in the inferential framework considered in our paper, the fully-blocked design can improve efficiency relative to these re-randomization designs. Also closely related to our paper is liu2022randomization, who extend the results in dasgupta2015causal to general stratified randomized designs. Their results specifically exclude the setting where each treatment is assigned exactly once per block, which is the primary setting that we consider in this paper.
The rest of the paper is organized as follows. In Section (ref) we describe our setup and notation. Section (ref) presents the main results. In Section (ref), we examine the finite sample behavior of various experimental designs via simulation in the context of $2^K$ factorial experiments. Finally, in Section (ref) we illustrate our proposed inference methods in an empirical application based on the experiment conducted in McKenzie2014. We conclude with recommendations for empirical practice in Section (ref).
Let $Y_i \in \mathbf R$ denote the observed outcome of interest for the $i$th unit. Let $D_i \in \mathcal D$ denote treatment status for the $i$th unit, where $\mathcal{D}$ denotes a finite set of values of the treatment. We assume $\mathcal D = \{1, \ldots, |\mathcal D|\}$. Generally, we use $D_i = 1$ to indicate the $i$th unit is untreated, but such a restriction is not necessary for our results. Let $X_i$ denote the observed baseline covariates for the $i$th unit, and denote its dimension by $\mathrm{dim}(X_i)$. For $d \in \mathcal{D}$, let $Y_i(d)$ denote the potential outcome for the $i$th unit if its treatment status were $d$. The observed outcome and potential outcomes are related to treatment status by the expression
We suppose our sample consists of $J_n := (|\mathcal{D}|)n$ i.i.d.\ units. For any random variable indexed by $i$, for example $D_i$, we denote by $D^{(n)}$ the random vector $(D_1, D_2, \ldots, D_{J_n})$. Let $P_n$ denote the distribution of the observed data $Z^{(n)}$ where $Z_i = (Y_i, D_i, X_i)$, and $Q_n$ denote the distribution of $W^{(n)}$, where $W_i = (Y_i(1), Y_i(2), \ldots, Y_i(|\mathcal{D}|), X_i)$. We assume that $W^{(n)}$ consists of $J_n$ i.i.d observations, so that $Q_n = Q^{J_n}$, where $Q$ is the marginal distribution of $W_i$. Given $Q_n$, $P_n$ is then determined by ((ref)) and the mechanism for determining treatment assignment. We thus state our assumptions in terms of assumptions on $Q$ and the treatment assignment mechanism.
Our object of interest will generically be defined as a vector of linear contrasts over the collection of expected potential outcomes across treatments. Formally, let \[\Gamma(Q) := (\Gamma_1(Q), \ldots, \Gamma_{|\mathcal{D}|}(Q))'~,\] where $\Gamma_d(Q) := E_Q[Y_i(d)]$ for $d \in \mathcal{D}$. Let $\nu$ be a real-valued $m\times |\mathcal{D}|$ matrix. We define \[\Delta_\nu(Q) := \nu\Gamma(Q) \in \mathbf R^m~,\] as our generic parameter of interest. For example, in the special case where $\mathcal{D} = \{1, 2\}$ and $\nu = (-1, 1)$, $\Delta_\nu(Q) = E_Q[Y_i(2) - Y_i(1)]$ corresponds to the familiar average treatment effect for a binary treatment. Further examples of $\Delta_\nu(Q)$ are provided in Examples (ref) and (ref) below.
We now describe our assumptions on $Q$. Our first assumption imposes restrictions on the (conditional) moments of the potential outcomes:
Assumption (ref)(a) is a mild restriction imposed to rule out degenerate situations and Assumption (ref)(b) is another mild restriction that permits the application of suitable laws of large numbers and central limit theorems. Assumption (ref)(c), on the other hand, is a smoothness requirement that ensures that units that are “close” in terms of their baseline covariates are also “close” in terms of their potential outcomes. Assumption (ref)(c) is a key assumption for establishing the asymptotic exactness of our proposed tests, since it allows us to argue that certain intermediate quantities in the derivations of our variance estimators vanish asymptotically (see for instance the proof of Lemma (ref)). Similar smoothness requirements are also imposed in bai2021inference.
Next, we specify our assumptions on the mechanism determining treatment status. In words, we consider treatment assignments which first stratify the experimental sample into $n$ blocks of size $|\mathcal{D}|$ using the observed baseline covariates $X^{(n)}$, and then assign one unit to each treatment uniformly at random within each block. We call such a design a matched tuples design. Formally, let \[ \lambda_j = \lambda_j(X^{(n)}) \subseteq \{1, \ldots, J_n\},~ 1 \leq j \leq n \] denote $n$ sets each consisting of $|\mathcal D|$ elements that form a partition of $\{1, \ldots, J_n\}$.
We assume treatment is assigned as follows:
We further require that the units in each block be “close” in terms of their baseline covariates in the following sense:
We will also sometimes require that the distances between units in adjacent blocks be “close" in terms of their baseline covariates:
We provide three examples of blocking algorithms which satisfy Assumptions (ref)--(ref):
In this section, we study estimation and inference for a general $m$-dimensional parameter $\Delta_\nu(Q)$ under a matched tuples design. For a pre-specified $\ell \times 1$ column vector $\Delta_0$ and $\ell \times m$ matrix $\Psi$ of rank $\ell$, the testing problem of interest is
at level $\alpha \in (0, 1)$. First we describe our estimator of $\Delta_\nu(Q)$. For $d \in \mathcal D$, define \[ \hat \Gamma_n(d) := \frac{1}{n} \sum_{1 \leq i \leq J_n} I \{D_i = d\} Y_i~, \] and let $\hat \Gamma_n = (\hat \Gamma_{n}(1), \ldots, \hat \Gamma_{n}(|\mathcal{D}|))'$. In words, $\hat\Gamma_n(d)$ is simply the sample mean of the observations with treatment status $D_i = d$, and $\hat \Gamma_n$ is the vector of sample means across all treatments $d \in \mathcal{D}$. With $\hat\Gamma_n$ in hand, our estimator of $\Delta_\nu(Q)$ is then given by \[\hat{\Delta}_{\nu,n} := \nu\hat\Gamma_n~.\] In what follows, it will be useful to define $\Gamma_d(X_i) := E[Y_i(d)|X_i]$. Our first result derives the limiting distribution of $\hat{\Delta}_{\nu, n}$ under our maintained assumptions.
To construct our test, we next define a consistent estimator for the asymptotic variance matrix $\mathbb V_\nu$. To begin, note by the law of total variance that \[ E[\operatorname*{Var}[Y_i(d) | X_i]] = \operatorname*{Var}[Y_i(d)] - E[E[Y_i(d) | X_i]^2] + E[Y_i(d)]^2~. \] Therefore, in order to estimate $\mathbb V_1$ consistently, it suffices to provide consistent estimators for $E[E[Y_i(d) | X_i]^2]$, $E[Y_i(d)]$, and $\operatorname*{Var}[Y_i(d)]$. A similar remark applies to $\mathbb V_2$. In light of this, define
To understand the construction, note that in order to estimate $E[E[Y_i(d) | X_i]^2]$ consistently, we would ideally average over the products of the outcomes of two units with similar values of $X_i$ and both with treatment status $d$. By construction, however, only one unit in each block has treatment status $d$. To overcome this problem, note that Assumption (ref) ensures that in the limit units in adjacent blocks also have similar values of $X_i$. Therefore, to construct our estimator of $E[E[Y_i(d)|X_i]^2]$, denoted by $\hat{\rho}_n(d,d)$, we average over the product of the outcomes of the units with treatment status $d$ in two adjacent blocks. $\hat \rho_n(d, d)$ is analogous to the “pairs of pairs" variance estimator in bai2021inference. A similar construction has also been used in abadie2008estimation in a related setting. On the other hand, for $d \neq d'$, we have distinct units with treatment status $d$ and $d'$ within each block, and therefore our estimator of $E[E[Y_i(d) | X_i] E[Y_i(d') | X_i]]$, denoted $\hat{\rho}_n(d,d')$, can be estimated using units within the same block.
Our estimator for $\mathbb V_\nu$ is then given by $\hat{\mathbb V}_{\nu, n} := \nu\hat{\mathbb V}_n\nu'$, where
with
Given this estimator, our test is given by \[\phi_n^{\nu}(Z^{(n)}) = I\{T_n^{\nu}(Z^{(n)}) > c_{1 - \alpha}\}~,\] where \[T_n^{\nu}(Z^{(n)}) = n(\Psi\hat\Delta_{\nu,n}- \Psi \Delta_0)'(\Psi\hat{\mathbb V}_{\nu, n}\Psi')^{-1}(\Psi\hat\Delta_{\nu,n} - \Psi \Delta_0)~,\] and $c_{1 - \alpha}$ is the $1 - \alpha$ quantile of the $\chi^2_\ell$ distribution. Our next result establishes the consistency of $\hat{\mathbb{V}}_n$ for $\mathbb{V}$ and the asymptotic validity of the above test.
Next, we study the properties of two commonly recommended inference procedures in the analysis of matched tuple designs. The first procedure is a $t$-test obtained from a linear regression of outcomes on treatment indicators and block fixed effects. Specifically, we consider a $t$-test obtained from the following regression:
which we interpret as the projection of $Y$ on the indicators for treatment status and block fixed effects. Let $\hat \beta_n(d)$, $d \in \mathcal D \backslash \{1\}$ and $\hat \delta_{j, n}$, $1 \leq j \leq n$ denote the OLS estimators of $\beta(d)$, $d \in \mathcal D \backslash \{1\}$ and $\delta_j$, $1 \leq j \leq n$. It is common in practice to use $\hat \beta_n(d)$ as an estimator for the pairwise average treatment effect between treatment $d$ and treatment $1$. See, for instance, McKenzie2013 and McKenzie2014. Furthermore, researchers often conduct inference on the pairwise ATEs using the heteroskedasticity-robust variance estimator obtained from (ref). Formally, for $d \in \mathcal D \backslash \{1\}$ and $\Delta_0 \in \mathbf R$, consider the problem of testing
at level $\alpha \in (0, 1)$. Let $\kappa_j\cdot \hat {\mathbb V}_n^{\rm sfe}(d, 1)$ denote the “HC$j$" heteroskedasticity-robust variance estimator of $\hat \beta_n(d)$ from the linear regression in (ref), where $\kappa_j$ for $j \in \{0, 1\}$ corresponds to one of two common degrees of freedom corrections mackinnon1985some: \[ \kappa_j =
\] The test is then defined as
where $z_{1 - \frac{\alpha}{2}}$ is the $(1 - \frac{\alpha}{2})$-th quantile of the standard normal distribution and
The following theorem shows that the OLS estimator $\hat \beta_n(d)$ is numerically equivalent to the standard difference-in-means estimator. However, it shows that the $t$-test defined in ((ref)) is not generally valid for testing the null hypothesis defined in ((ref)).
bai2021inference remark that the test defined in (ref) is conservative in the context of a matched-pair design when using $\mathrm{HC}1$. Theorem (ref) shows that, when considering a matched tuples design with more than two treatments, this is no longer necessarily the case.
The second procedure is a block-cluster robust $t$-test which modifies a recent proposal in chaisemartin2022at to the setting with multiple treatments. Specifically, we consider a cluster-robust $t$-test constructed from a regression of outcomes on a constant and treatment indicators: \[ Y_i = \gamma(1) + \sum_{d \in \mathcal D \backslash \{1\}} \gamma(d) I \{D_i = d\} + \epsilon_i~, \] where clusters are defined at the level of blocks of units $\{\lambda_j\}_{1 \le j \le \mathcal{D}}$. Let $\hat{\gamma}_n(d)$, $d \in \mathcal{D} \backslash \{1\}$ denote the OLS estimator of $\gamma(d)$, it then follows immediately that $\hat{\gamma}_n(d) = \hat{\Gamma}_n(d) - \hat{\Gamma}_n(1)$. We then consider the problem of testing (ref) at level $\alpha \in (0, 1)$ using a test defined by \[\phi^{\rm bcve}_n(Z^{(n)}) = I\{|T^{\rm bcve}_n(Z^{(n)})| > z_{1 - \frac{\alpha}{2}}\}~,\] where $z_{1 - \frac{\alpha}{2}}$ is the $(1 - \frac{\alpha}{2})$-th quantile of the standard normal distribution and
with $\hat {\mathbb V}_n^{\rm bcve}(d)$ denoting the $d$-th diagonal element of the block-cluster variance estimator defined as:
where $C_i = (1, I \{D_i = 2\},\dots, I \{D_i = |\mathcal{D}|\})'$ and $\hat \epsilon_i = \sum_{d \in \mathcal{D}\backslash \{1\}} (Y_i - \hat \gamma_n(d)) I\{D_i = d\} + Y_i I\{D_i=1\}- \hat\gamma_n(1)$.
The following theorem shows that the $t$-test defined in (ref) is generally conservative for testing the null hypothesis defined in (ref).
Our analysis so far has focused on the setting where $J_n = |\mathcal{D}| n$ units are blocked into $n$ blocks of size $|\mathcal{D}|$, and each treatment $d \in \mathcal{D}$ is assigned exactly once in each block. In this section, we consider a modification of this design where units are grouped into blocks of size $2|\mathcal{D}|$ and each treatment status $d \in \mathcal{D}$ is assigned exactly twice in each block. Formally, for the remainder of this section suppose $n$ is even, and let \[ \tilde \lambda_j = \tilde \lambda_j(X^{(n)}) \subseteq \{1, \ldots, J_n\},~ 1 \leq j \leq n / 2\] denote $n/2$ sets each consisting of $2 |\mathcal D|$ elements that form a partition of $\{1, \ldots, J_n\}$.
We assume treatment is assigned as follows:
We further require that the units in each block be “close” in terms of their baseline covariates in the following sense:
We first establish that the limiting distribution of $\hat{\Delta}_{\nu,n}$ for such a “replicate” design is the same as that for the matched tuples design considered in Theorem (ref).
Although the limiting distribution of $\hat{\Delta}_{\nu,n}$ for the standard matched tuples and replicate designs are identical, variance estimation in the replicate design is often understood to be conceptually simpler, because each treatment status is assigned twice in each block athey2017econometrics. Indeed, in this case an alternative variance estimator can be constructed which is identical to the estimator proposed in Section (ref) except that we replace $\hat \rho_n(d,d)$ by
which no longer requires averaging over the product of outcomes of units in adjacent blocks. The following theorem establishes the consistency of $\tilde \rho_n(d, d)$, where importantly we note that Assumption (ref), which maintains that adjacent blocks be “close", is no longer required. It is then straightforward to show the consistency of the corresponding variance estimator for $\hat \Delta_{\nu, n}$ constructed by replacing $\hat{\rho}_n(d,d)$ in $\hat{\mathbb V}_n$ with $\tilde \rho_n(d, d)$.
We remark that Theorems (ref)--(ref) and Theorems (ref)--(ref), yielding identical conclusions, do not allow us to effectively compare the properties of the standard matched tuples design and matched tuples with replicates. In order to compare these designs, we evaluate their finite sample properties via simulation in Section (ref). There, we find that the mean squared error of $\hat{\Delta}_{\nu,n}$ under the replicate design is typically larger than under the standard non-replicate design. However, we also find that the rejection probabilities of our proposed tests under the replicate design are much closer to the nominal level relative to the non-replicate design, which can sometimes exhibit rejection probabilities strictly smaller than the nominal level when matching on multiple covariates. As a result, the replicate design is sometimes able to achieve better power relative to the non-replicate design. We emphasize, however, that our current asymptotic framework is not precise enough to capture these differences. One possible conjecture is that since replicate designs could be thought of as convex combinations of matched tuples designs bai2022optimality, it is as if we are averaging over multiple matched tuples designs when we estimate the limiting variance. However, we leave a detailed theoretical comparison of these two designs to future work.
In this section we apply the results derived in Sections (ref)--(ref) to study the asymptotic properties of what we call “fully-blocked" $2^K$ factorial designs. Section (ref) introduces $2^K$ factorial experiments. Section (ref) introduces the fully-blocked factorial design and compares the efficiency properties of fully-blocked factorial designs to some alternative designs.
In this section we describe the setup of a $2^K$ factorial experiment, the resulting parameters of interest, and their corresponding estimators wu2011experiments. A $2^K$ factorial design assigns treatments which are combinations of multiple “factors," where each factor can take two distinct values, or “levels.” For instance, karlan2014agricultural study the effect of capital constraints and uninsured risk on the investment decisions of farmers in Ghana. In their setting, each treatment consists of two factors: whether or not a household receives a cash grant, and whether or not a household receives an insurance grant. Our setup and notation mirror the framework introduced in dasgupta2015causal and li2020rerandomization. Given $K$ factors each with two treatment levels $\{-1, +1\}$, our set of treatments $\mathcal{D}$ now consists of all possible $2^K$ factor combinations. For a factor combination $d \in \mathcal{D}$, define $\iota_k(d) \in \{-1, +1\}$ to be the level of factor $k$ under treatment $d$. The vector $\iota(d) := (\iota_1(d), \iota_2(d), \ldots, \iota_K(d))$ then describes the levels of all $K$ factors associated with factor combination $d$. This notation allows us to define factorial effects as parameters of the form $\Delta_{\nu}(Q)$ for appropriately constructed contrast vectors $\nu$. For instance, consider the contrast vector defined as \[\nu_k := \left(\iota_k(1), \iota_k(2), \ldots, \iota_k(|\mathcal{D}|)\right)~.\] Then, the parameter $\Delta_{\nu_k}(Q)$ obtained from this contrast can be written as
We define the main effect of factor $k$ as $2^{-(K-1)}\Delta_{\nu_k}(Q)$. In words, the main effect of factor $k$ measures the average difference between the outcomes of factor combinations under which the $k$th factorial effect is $1$ versus the outcomes of factor combinations under which the $k$th factorial effect is $-1$. The re-scaling $2^{-(K-1)}$ is introduced because there are $2^{K-1}$ possible values for all the factor combinations when fixing the $k$th factor. We call $\nu_k$ the generating vector for the main effect of factor $k$.
We can subsequently build on the generating vectors of the main effects in order to define the interaction effects between various factors. The interaction effect between a given set of factors is defined using the contrast obtained from taking the element-wise product of the generating vectors for the relevant factors. For instance, the two-factor interaction between factors $k$ and $k'$ is defined as $2^{-(K-1)}\Delta_{\nu_{k,k'}}(Q)$, where $\nu_{k,k'} := \nu_k \odot \nu_{k'}$ and $\odot$ denotes element-wise multiplication. Similarly, the three-factor interaction $2^{-(K-1)}\Delta_{\nu_{k,k',k''}}(Q)$ is defined using the contrast vector $\nu_{k,k',k''} := \nu_k \odot \nu_{k'} \odot \nu_{k''}$. We illustrate these definitions in the special case of a $2^2$ factorial design in Example (ref) below. For simplicity, in what follows, we omit the re-scaling by $2^{-(K-1)}$ in our discussions and results.
Given the above setup, we estimate the factorial effect given by $\Delta_{\nu}(Q)$ using the estimator $\hat{\Delta}_{\nu,n}$ defined in Section (ref). wu2011experiments and dasgupta2015causal explain that $\hat{\Delta}_{\nu,n}$ is a standard estimator in this context. For instance, the estimator of the main effect of factor $k$, $2^{-(K-1)}\hat \Delta_{\nu_k,n}$, is in fact the difference-in-means estimator over the $k$-th factor:
In this section, we compare the asymptotic variance of the estimator $\hat\Delta_{\nu,n}$ under what we call a “fully-blocked" factorial design relative to some alternative designs. A fully-blocked factorial design first blocks the experimental sample into $n$ blocks of size $2^K$ based on the observable characteristics $X^{(n)}$, and then assigns each of the $2^K$ factor combinations exactly once in each block. Formally, a fully-blocked factorial design is simply a matched tuples design as defined in Section (ref), where $\mathcal{D}$ consists of the set of all possible factor combinations.
Our first result compares the fully-blocked factorial design to completely randomized and stratified factorial designs. Given a $2^K$ factorial experiment and a sample of size $J_n = n2^K$, a completely randomized factorial design simply assigns $n$ individuals to each of the $2^K$ factor combinations at random. A stratified factorial design first partitions the covariate space into a finite number of groups, or “strata", and then performs a completely randomized factorial design within each stratum. Formally, let $h: \mathrm{supp}(X) \to \{1, \ldots, S\}$ be a function which maps covariate values into a set of discrete strata labels. Then, a stratified factorial design performs a completely randomized factorial design within each stratum produced by $h(\cdot)$. Note that a completely randomized design is a special case of the stratified factorial design where the co-domain of $h(\cdot)$ is a singleton. See rubin2016 and li2020rerandomization for further discussion of these designs. Theorem (ref) shows that the asymptotic variance of $\hat{\Delta}_{\nu,n}$ is weakly smaller under a fully-blocked factorial design than that under any stratified factorial design as defined above, as long as the potential outcomes satisfy the smoothness assumptions described in Assumption (ref)(c).
Our next result considers settings where only a subset of the factors are of primary interest to the researcher. For instance, besedevs2012age use a factorial design to study how the number of options in an agent's choice set affects their ability to make optimal decisions. Here the primary factor of interest is the number of options (four or thirteen), but the design also features other secondary factors. In such a case we might imagine that a matched pairs design which focuses on the factor of primary interest and assigns the other factors by i.i.d.\ coin flips may be more efficient for estimating the primary factorial effect than the fully-blocked design which treats all the factors symmetrically. In particular, we consider a setting where we are interested in the average main effect on the $k$th factor, $\Delta_{\nu_k}(Q)$, and compare the performance of the fully-blocked design to a design which performs matched pairs over the $k$th factor while assigning the other factors to individuals at random using i.i.d.\ Bernoulli(1/2) assignment. We call such a design the “factor $k$ specific" matched pairs design. Formally, let \[ \zeta_j = \zeta_j(X^{(n)}) \subset \{1, \dots, 2^K n\},~ 1 \leq j \leq 2^{K - 1} n \] denote a partition of the set of indices such that each $\zeta_j$ contains two units. The “factor $k$ specific" matched pairs design satisfies the following assumption:
Theorem (ref) shows that the asymptotic variance of $\hat{\Delta}_{\nu_1,n}$ is weakly smaller under a fully-blocked design than that under the factor specific matched pairs design.
In this section we examine the finite sample performance of the estimator $\hat\Delta_{\nu,n}$ and the test $\phi^\nu_n(Z^{(n)})$ in the context of a $2^K$ factorial experiment, under various alternative experimental designs. In Sections (ref) and (ref) the data generating processes are as specified below (in Section (ref) we study an alternative design with multiple covariates and factors). For $d = (d^{(1)}, d^{(2)}) \in \{-1, 1\}^2$ and $1\leq i \leq 4 n$, the potential outcomes are generated according to the equation:
In each of the specifications, $((X_i, \epsilon_{i}): 1\leq i \leq 4 n)$ are i.i.d; for $1 \leq i \leq 4n$, $X_i$ and $\epsilon_{i}$ are independent.
We consider five parameters of interest as listed in Table (ref). $\Delta_{\nu_1}(Q)$ and $\Delta_{\nu_2}(Q)$ correspond to the main factorial effects for the two factors. $\Delta_{\nu_{1,2}}(Q)$ corresponds to the interaction effect between the two factors, as discussed in Example (ref). $\Delta_{\nu_1^1}(Q)$ and $\Delta_{\nu_{-1}^1}(Q)$ denote the average effect of one factor, keeping the value of the other factor fixed at $1$ or $-1$. All simulations are performed with a sample of size $4n = 1000$.
In this section, we study the mean-squared-error performance of $\hat\Delta_{\nu,n}$ across several experimental designs. We analyze and compare the MSE for all five parameters of interest for the following seven experimental designs:
Table (ref) displays the ratio of the MSE of each design relative to the MSE of MT, computed across 4,000 Monte Carlo replications. In each of the designs, we set treatment effects to zero by setting $\tau=0$. As expected from Theorems (ref) and (ref), MT outperforms B-B, C, MP-B, Large-2, and \textbf{Large-4} in every model specification. We also find that \textbf{MT} compares favorably to \textbf{RE}, with \textbf{RE} slightly outperforming \textbf{MT} in some cases, but with \textbf{MT} outperforming in general. Although we do not have formal results comparing the matched tuples design to re-randomization, we note that re-randomization redraws treatments until the distances between certain features of the covariate distribution across treatment statuses are below certain pre-specified thresholds. In contrast, the matched tuples design attempts to \emph{minimize} these distances by blocking units finely based on the covariates. See also Remark 3 of bai2022optimality for a related observation in the binary treatment setting.
In this section, we study the finite sample properties of several different tests of the null hypothesis $H_0: \Delta_{\nu} = 0$ for various choices of $\nu$, against the alternative hypotheses implied by setting $\tau = 0.2$. In this section we restrict our attention to five assignment mechanisms: B-B, C, MT, Large-2 and Large-4. We exclude MP-B because it is a non-standard experimental design for which we have not developed an inference procedure. We also exclude the re-randomization design (\textbf{RE}) because, although it is a widely studied design, the inferential results in li2020rerandomization are derived in a finite population framework which is distinct from our super-population framework, and their resulting limiting distribution is non-normal.
In each case we perform our hypothesis tests at a significance level of $0.05$. For design B-B, tests are performed using a standard $t$-test. For designs C, Large-2 and Large-4 the tests are constructed using the asymptotic normality result from Theorem (ref) combined with variance estimators constructed using the same plug-in method as in bugni2018 and bugni2019inference. For design {\bf MT} the test is constructed as described in Theorem (ref). Table (ref) displays the rejection probabilities under the null and alternative hypotheses, computed from 2,000 Monte Carlo replications. The results show that the rejection probabilities are universally around 0.05 under the null hypothesis, which verifies the validity of our tests across all the designs. Under the alternative hypotheses implied by $\tau=0.2$, the rejection probabilities vary substantially across the different designs, outcome models and parameters. However, our matched tuples design displays the highest power for almost all parameters and model specifications.
In this section we repeat the previous simulation exercises while varying the number of factors $K$ and the number of observed covariates $\mathrm{dim}(X_i)$. The data generating process is constructed as follows:
where $\tau \in \{0, 0.1\}$, $d = (d^{(1)}, \ldots, d^{(K)})$ and $d^{(k)}\in \{-1, 1\}$ represents the treatment status of the $k$-th factor. We set $\gamma_{d} = 1$ if $d^{(2)}=1$, $\gamma_{d} = -1$ otherwise, in order to ensure the conditional means are heterogeneous in the second factor. $\tilde{X}_i$ contains $9$ covariates, out of which the first $\text{dim}(X_i)$ covariates are observed and used for the experimental designs. The distributions of $\tilde{X}_i, \epsilon_i$ and the values of $\beta$ are calibrated using data obtained from rubin2016, who study the covariate balancing properties of $2^K$ factorial re-randomization designs using data from the New York Department of Education (NYDE). Details on the empirical context and construction of the data generating process are provided in Appendix (ref).
To construct our matched tuples of size $2^K$ when $\mathrm{dim}(X_i) > 1$, we employ the recursive pairing algorithm described in Section (ref) using the Mahalanobis distance. We emphasize, however, that this approach is not guaranteed to be optimal, and we leave the study of potentially more effective matching algorithms to future work.
In addition to the standard matched tuples design (MT), we also include a matched tuples design with a replicate for each treatment as described in Section (ref), denoted by MT2. For example, in the MT2 design with two factors, units are matched into groups of eight, and two units receive each factor combination. We also continue to consider the alternative designs (C, \textbf{Large-4}, \textbf{MP-B} and \textbf{RE}) from Section (ref). When constructing the strata for \textbf{Large-4}, we stratify on one covariate drawn at random from the set of available covariates.
In Table (ref) we report the ratio of the MSE of each design relative to the MSE of {\bf MT} when $\text{dim}(X_i) = 1$ and $K = 1$ (computed from 4,000 Monte Carlo replications). For all experiments in this section, the number of observations is fixed to be 1,280 so that we have 20 matched tuples of size 64 when $K=6$. Our simulation results are consistent with those in Section (ref): MT displays the lowest MSE across almost all model specifications. Although MT2 generally produces larger MSEs than MT, it still performs favorably relative to the other designs. For methods that use an increasing number of covariates when $\mathrm{dim}(X_i)$ increases (MT, MT2, MP-B and \textbf{RE}), we observe that the MSE in fact \emph{increases} with the number of available covariates. We expect this is because (as shown in Appendix (ref)) the first covariate is a much stronger predictor of the control outcome than the other available covariates, which are relatively uninformative.
In Table (ref), we compute the rejection probabilities when testing the null hypothesis $H_0: \Delta_{\nu_1} = 0$ against the alternative implied by setting $\tau = 0.1$, for various choices of $K$ and $\mathrm{dim}(X_i)$ (computed from 1,000 Monte Carlo replications). Under the null hypothesis, we observe that our tests under design MT become conservative as $\mathrm{dim}(X_i)$ and $K$ increase. In particular, we notice a large difference between $K = 4$ and $K = 5$. However, despite being conservative, MT still displays favorable power properties relative to C and Large-4 for all but the largest choices of $K$.
Our next observation is that our tests under design MT2 remain exact even as $\mathrm{dim}(X_i)$ and $K$ both increase. As we explain in Section (ref), we suspect that our challenges for inference using MT come from poor estimation of the variance, which seems to be alleviated in MT2, where the number of observations receiving each treatment within a tuple are doubled. As a result of this exactness, MT2 achieves higher power than MT when $\mathrm{dim}(X_i)$ and $K$ are large. To further explore these power improvements, Figure (ref) presents power plots for three specific choices of $K$ and $\mathrm{dim}(X_i)$ with $\tau$ ranging from 0 to 0.1 (Figure (ref) in the appendix presents power plots for alternatives implied by larger values than $\tau = 0.1$). First, when $\mathrm{dim}(X_i)$ and $K$ are small, for instance $\mathrm{dim}(X_i)=K=1$, we observe no significant difference between the power plots generated by MT and \textbf{MT2}. However, when the dimension of the covariates and factors are both large, for instance $\mathrm{dim}(X_i)=6, K=4$, \textbf{MT2} dominates \textbf{MT} for all alternative hypotheses. Therefore, our recommendation to practitioners is to consider a matched tuples design when working with few treatments and covariates, but to consider the replicated design when dealing with a large number of treatments and/or covariates.
In this section, we illustrate the inference procedures introduced in Section (ref) using the data collected in McKenzie2014\footnote{The original paper features six rounds of surveys which were pooled in the final analysis. We perform our analysis exclusively on the data obtained in the sixth round in order to avoid complications related to time-series dependence across rounds. For simplicity, we additionally drop quadruplets with missing values, and 4 “leftover” groups whose sizes range from 5 to 8 firms. This results in a final sample of 120 quadruplets, or $4n = 480$. Further results on the long-run effects (collected in a seventh survey wave) are contained in Table (ref) in Section (ref) of the appendix.}. McKenzie2014 conduct a randomized experiment in order to investigate the effects of several capital aid programs on the profits of small businesses in Ghana. In their experiment, there are three treatment arms, where (in our notation) $D_i = 1$ indicates that the $i$th firm is untreated, $D_i = 2$ indicates being offered cash, and $D_i = 3$ indicates being offered in-kind grants. The null hypotheses of interest are
for $d \in \{2, 3\}$, as well as
In their experimental design, blocks are defined by quadruplets, where each quadruplet contains two untreated units with $D_i = 1$, one treated unit with $D_i = 2$, and one treated unit with $D_i = 3$. Despite the slight departure from the framework presented in Sections (ref)--(ref), in that there are two untreated units in each quadruplet, we show in Appendix (ref) that a slight modification of the variance estimator in Theorem (ref) produces a valid test for (ref)--(ref). Specifically, we pretend that there are four treatment levels in each quadruplet, while the first two are in fact controls. Then, by setting generating vectors $\nu^{2}=(-1/2,-1/2,1,0)$, $\nu^{3} = (-1/2, -1/2, 0, 1)$, and $\nu^{2,3}=(0,0,-1,1)$ and proceeding with the testing procedure in Theorem (ref), we obtain valid tests for $H_0^d$ and $H_0^{2,3}$. For each of the hypotheses in (ref)--(ref), we implement the following tests:
We note that McKenzie2014 test (ref) and (ref) using a $t$-test obtained from a linear regression of outcomes on treatment indicators and block fixed effects. However, as was shown in Theorem (ref), such a procedure is not guaranteed to be valid. On the other hand, we expect that the $t$-test obtained from a linear regression without block fixed effects should be conservative for testing (ref)--(ref) in light of the observations made in Example (ref) and the fact that this test coincides with a standard two-sample $t$-test.
Our results are presented in Table (ref). The point estimates of the two methods are identical because the OLS estimator coincides with the difference-in-means estimator. However, the standard errors obtained from our variance estimator are always smaller than the heteroskedasticy-robust standard errors. For example, when testing (ref) for $d = 3$ among the female subsample, the standard error produced from our variance estimator is 15.21 whereas the heteroskedasticy robust standard error is 18.13. We note that overall the improvements are modest; this suggests that the conditional expectation of the outcomes does not vary substantially with the observable characteristics in this survey wave. This is further corroborated by the calibrated simulations presented in Table (ref) in Appendix (ref).
We conclude with some recommendations for empirical practice based on our theoretical results as well as the simulation study above. For inference about the linear contrast of expected outcomes given by $\Delta_{\nu}$ in a matched tuples design, we recommend the test $\phi_n^{\nu}$ defined in Section (ref): our simulations results show that this test does a good job of controlling size in large samples (approximately 80 blocks). We have shown that tests based on the heteroskedasticity-robust variance estimator from a linear regression of outcomes on treatment and block fixed effects may be invalid, in the sense of having rejection probability strictly greater than the nominal level under the null hypothesis. Tests based on the heteroskedasticity-robust variance or block-cluster variance estimators from a linear regression of outcomes on treatment are valid but potentially conservative, which would result in a loss of power relative to our proposed test.
We also find that matched tuples designs have favorable efficiency properties relative to other popular designs (with a specific illustration in the setting of $2^K$ factorial designs). However, this comes with the caveat that when dealing with a large number of treatments (in our simulations, this translated to having fewer than 80 blocks) and/or large number of covariates, practitioners may want to consider the replicated matched tuples design introduced in Section (ref), as our simulations suggest that this design may have more robust size control, which translates to better power in such cases.