EconBase
← Back to paper

Matching Estimators with Few Treated and Many Control Observations

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,072 characters · 10 sections · 53 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Matching Estimators with Few Treated and Many Control Observations

\def\spacingset#1{ {#1}} \spacingset{1}

\newsavebox{\tablebox} \newlength{\tableboxwidth}

center[center omitted — 148 chars of source]

We analyze the properties of matching estimators when there are few treated, but many control observations. We show that, under standard assumptions, the nearest neighbor matching estimator for the average treatment effect on the treated is asymptotically unbiased in this framework. However, when the number of treated observations is fixed, the estimator is not consistent, and it is generally not asymptotically normal. Since standard inference methods are inadequate, we propose alternative inference methods, based on the theory of randomization tests under approximate symmetry, that are asymptotically valid in this framework. We show that these tests are valid under relatively strong assumptions when the number of treated observations is fixed, and under weaker assumptions when the number of treated observations increases, but at a lower rate relative to the number of control observations.

\

{\it Keywords:} matching estimators, treatment effects, hypothesis testing, randomization inference, synthetic control estimator

\

{\it JEL Codes:} C12; C13; C21

\spacingset{1.45}

\onehalfspacing

Introduction

Matching estimators have been widely used for the estimation of treatment effects under a conditional independence assumption (CIA).\footnote{See, for example, Imbens2004, Imbens_Wooldrige, and Imbens_examples for reviews.} In many cases, matching estimators have been applied in settings where (1) the interest is on the average treatment effect for the treated (ATT), and (2) there is a large reservoir of potential controls (see Imbens_Wooldrige). AI_2006 (henceforth, AI) study the asymptotic properties of nearest-neighbor (NN) matching estimators when the number of control observations ($N_0$) grows at a faster rate than the number of treated observations ($N_1$). However, their asymptotic theory still depends on both the number of treated and control observations going to infinity. Therefore, reliance on such asymptotic approximation should be considered with caution when the number of treated observations is small, even if the total number of observations is large.

In this paper, we analyze the properties of NN matching estimators when $N_1$ is fixed, while $N_0$ goes to infinity. We first show that the NN matching estimator is asymptotically unbiased for the ATT, under standard assumptions used in the literature on estimation of treatment effects under selection on observables.\footnote{This is true whether asymptotic unbiasedness is defined based on the limit of the expected value of the estimator, or based on the expected value of the asymptotic distribution.} This is consistent with the conclusions from AI, who show that the conditional bias of the NN matching estimator can be ignored, provided that $N_0$ increases fast enough, relative to $N_1$. In their setting, the NN matching estimator is consistent and asymptotically normal. In our setting, however, the variance of the estimator does not converge to zero, and the estimator will not generally be asymptotically normal.\footnote{Our setting is different from the case of limited overlap considered by KT. The problem we analyze can arise even when the overlap condition considered by AI in their Assumption 2$'$(ii) is satisfied. The difference relative to the case considered by AI is that $N_1$ remains fixed, so it is not possible to apply a law of large numbers and a central limit theorem on the average of the errors of the treated observations. } Our theory complements the theory developed by AI, providing a better approximation to settings in which there is a larger number of control relative to treated observations, but $N_1$ is not large enough, so that we cannot rely on asymptotic results in which $N_1$ goes to infinity.\footnote{The finite sample properties of matching and other related estimators have been evaluated in simulations by, for example, Frolich, Busso, Huber, and Bodory. In contrast to their approach, we provide theoretical and simulation results holding the number of treated observations fixed, but relying on the number of control observations going to infinity.}

The fact that the NN matching estimator is not asymptotically normal, in our setting, poses important challenges when it comes to inference. Inference based on the asymptotic distribution of the matching estimator derived by AI should not provide a good approximation when $N_1$ is very small, even if there are many control observations. The bootstrap procedure proposed by Otsu also relies on the number of both treated and control observations going to infinity. Leung consider a finite population setting with limited overlap, where the probability of treatment may converge to zero for some strata. While they provide conditions in which standard inference methods remain asymptotically valid in this case, our setting with $N_1$ fixed would not satisfy their conditions. Rothe provides robust confidence intervals for average treatment effects under limited overlap. For the case with continuous covariates, he combines his method with subclassification on the propensity score. However, with few treated and many control observations, it would not be possible to reliably estimate a propensity score. Moreover, while armstrong2017finitesample propose confidence intervals that are asymptotically valid even when $\sqrt{n}-$inference is not possible, the conditions they consider for this result are not satisfied in our setting.\footnote{armstrong2017finitesample present finite-sample results considering a setting in which errors are normal with known variance. Then they relax these conditions and consider a feasible version of their confidence intervals that is asymptotically valid. However, with $N_1$ fixed, we cannot have condition (21) in their paper being satisfied. Therefore, the results from their Theorem 4.2 cannot be directly applied to our setting. } Finally, for finite samples, rosenbaum1984 and rosenbaum2002 consider permutation tests for observational studies under strong ignorability. However, these tests rely on restrictive assumptions.\footnote{rosenbaum1984 assumes that the propensity score follows a logit model, while rosenbaum2002 assumes that observations are matched in pairs such that the probability of treatment assignment is the same conditional on the pair. }

Given the limitations of existing inference methods for the setting we analyze, we consider alternative inference methods based on the theory of randomization tests under an approximate symmetry assumption, developed by Canay. We focus on a test based on sign changes. We show that, under symmetry assumptions on the errors and on the heterogeneous treatment effects, this test provides asymptotically valid hypothesis testing for the ATT when $N_0 \rightarrow \infty$, even when $N_1$ is fixed. When $N_1$ increases, but at a lower rate than $N_0$, we show that this test is asymptotically valid even when we relax such symmetry conditions. Therefore, this test works with very few treated observations under relatively strong assumptions, and with a larger number of treated observations under weaker assumptions. We consider in Appendix (ref) an alternative test based on permutations, which also has the property of being valid under stronger assumptions when $N_1$ is fixed, and under weaker assumptions when $N_1$ increases.

The remainder of this paper proceeds as follows. We present our theoretical setup in Section (ref). In Section (ref), we derive the asymptotic distribution of the NN matching estimator, and derive conditions under which it is asymptotically unbiased in a setting with fixed $N_1$. In Section (ref), we consider an alternative inference method that is asymptotically valid when $N_0 \rightarrow \infty$, while $N_1$ remains fixed. We also consider the properties of this test when $N_1$ increases. In Section (ref), we present Monte Carlo (MC) simulations. In Section (ref), we contrast the different inference procedures in light of the theoretical results presented in Section (ref) and the simulations presented in Section (ref), providing guidance on which method should be chosen depending on the setting. We present in Section (ref) an empirical illustration based on the “Jovem de Futuro” program, which provides an example in which matching estimators could be used in settings with few treated and many control observations. Concluding remarks, including a discussion on the implications of our results for other types of matching estimators and for Synthetic Control applications, are presented in Section (ref).

Setting and Notation

We are interested in estimating the effect of a binary treatment ($W$) on some outcome ($Y$). Following Rubin1973, we define $Y(0)$ as the potential outcome under no exposure to treatment, and $Y(1)$ as the potential outcome under exposure to treatment. Therefore, the observed outcome is given by $Y = W Y(1) + (1-W)Y(0)$. In addition to $Y$ and $W$, we also consider a continuous random vector of $k$ real-valued pretreatment variables, which we denote by $X$.\footnote{We discuss in Appendix (ref) cases in which components of $X$ are discrete, and cases in which components of $X$ have a mixed distribution.}

We observe $N_1$ treated observations obtained by random sampling from the distribution of $(Y,X)|W=1$ and $N_0$ untreated observations obtained by random sampling from the distribution of $(Y,X)|W=0$. Let $\mathcal{I}_w$ denote the set of indexes for observations with $W_i=w$.

assumption[Sample] $\{Y_i , X_i , W_i \}_{i \in \mathcal{I}_1 \cup \mathcal{I}_0}$ is a pooled sample of $N_1$ treated ($i \in \mathcal{I}_1$) and $N_0$ untreated ($i \in \mathcal{I}_0$) observations obtained by random sampling from their respective population counterparts. Furthermore, observations in the treated and control samples are independent.

We consider the case in which $N_1$ is fixed, while $N_0$ goes to infinity. One possibility is that there is a large set of units that could potentially be treated, but only a finite number of them actually receive treatment. For example, in the empirical application, to be presented in Section (ref), there is a large number of schools that could potentially receive the treatment, but only a small number of them actually received it. Alternatively, we can imagine that there is a large number of treated units, but we only have data from a small sample of them. Assumption (ref) is similar to Assumption 3$'$ from AI and from the first condition stated in Theorem 1 from AI_martingale, in that the proportions of treated and control observations in the sample may not reflect their proportions in the population.

The goal is estimating the ATT, which we denote by

eqnarray[eqnarray omitted — 73 chars of source]

We focus on an estimand related to the treatment effect on the treated because, given our setting with $N_1$ finite and $N_0$ large, we would only have a small number of treated observations to serve as potential neighbors to estimate the counterfactual of the control observations, in case we wanted to estimate the average treatment effect (ATE). Likewise, AI consider the estimation of the ATT when they consider an asymptotic framework in which $N_0$ grows at a faster rate than $N_1$. In Appendix (ref) we discuss the case in which the estimand of interest is the ATT conditional on the realization of the covariates for the treated observations, $\{X_i \}_{i \in \mathcal{I}_1}$.

Assumption (ref) does not impose any restriction on how the distribution of $(Y(1),Y(0),X)$ conditional on $W = w$ depends on $w$. The following assumption restricts the way in which these distributions may differ, which is a standard conditional independence assumption (CIA).

assumption[Conditional Independence Assumption] $Y(0) \perp\!\!\!\perp W | X$.

While Assumption (ref) restricts that the conditional distribution of $Y(0)$ given $X$ is the same for both treatment and control observations, the density of $X$ conditional on $W=w \in \{0,1\}$ can potentially depend on $w$. This is what potentially generates bias in a simple comparison of means between treated and control groups, without taking into account that these groups might have different distributions of covariates $X$. We do not need to impose conditional independence of $Y(1)$ because the focus is on the ATT, and not on the average treatment effects.

The next assumption states conditions on the distribution of the covariates. Let $\mathbb{X}_w$ be the support of $X$ conditional on $W=w$, and $f_w : \mathbb{X}_w \rightarrow \mathbb{R}_+$ be the conditional density of $X$ given $W=w$, for $w \in \{0,1\}$.

assumption[Distribution of covariates] (i) $X \in \mathbb{R}^{k}$ is an absolutely continuous random vector, (ii) $\mathbb{X}_1 \subseteq \mathbb{X}_0$, where $\mathbb{X}_1$ and $\mathbb{X}_0$ are compact, (iii) $f_1$ and $f_0$ are differentiable for all points in the interior of their support, bounded from above in $\mathbb{X}_1$, and $f_0$ is bounded from below in $\mathbb{X}_1$, and (iv) for all points in $x \in \mathbb{X}_1$ at least a fraction $\phi$ of any sphere around $x$ belongs to $\mathbb{X}_1$.

This assumption guarantees that, for each $i$ in the treated group, we can find an observation $j$ in the control group with covariates $X_j$ arbitrarily close to $X_i$ when $N_0 \rightarrow \infty$. As we show in Appendix (ref), Assumption (ref) implies that there is an $\eta > 0$ such that, for all $x \in \mathbb{X}_1$, $Pr(W = 1 | X=x) < 1-\eta$.\footnote{This propensity score is defined over the distribution of $(Y,X,W)$. }

The main identification problem arises from the fact that we observe either $Y_i(1)$ or $Y_i(0)$ for each observation $i$. If we had two observations, $i \in \mathcal{I}_1$ and $j \in \mathcal{I}_0$, with $X_i=X_j=x$, then, under Assumptions (ref) and (ref), $\mathbb{E}[Y_i | X_i=x] - \mathbb{E}[Y_j | X_j=x] = \mathbb{E}[Y(1) | W=1,X=x] - \mathbb{E}[Y(0) | W=0,X=x]=\mathbb{E}[Y(1) - Y(0)| X=x,W=1]$. The main challenge is that, with a continuous random variable $X$, the probability of finding treated and control observations with exactly the same $X$ is zero. The idea of the NN matching estimator is to input the missing potential outcome of a treated observation $i \in \mathcal{I}_1$ with observations from the control group $j \in \mathcal{I}_0$ that are as close as possible in terms of covariates $X_i$. More specifically, for a distance metric $d(a,b)$ in $\mathbb{R}^k$, let $\mathcal{J}_M(i)$ be the set of $M$ nearest neighbors in the control group of observation $i \in \mathcal{I}_1$. Then the NN matching estimator is given by

eqnarray[eqnarray omitted — 136 chars of source]

where we consider the matching estimator with replacement. We consider the case in which $d(a,b) = [(a-b)'V(a-b)]^{1/2}$ for some positive definite matrix $V$. In Remark (ref) in the appendix we show that our results are also valid if we consider the Mahalanobis distance.

Asymptotic Unbiasedness and Asymptotic Distribution

For $w \in \{ 0,1\}$, we define $\mu(x,w) = \mathbb{E}[Y| X=x,W=w]$ and $\epsilon = Y - \mu(X,W)$. Since we are focusing on the average treatment effect on the treated, we also define $\mu_w(x) = \mathbb{E}[Y(w)| X=x,W=1]$.\footnote{AI define $\mu_w(x) = \mathbb{E}[Y(w)| X=x]$. We use a slightly different definition because we focus on the ATT. } Under Assumption (ref), we have that $\mu(x,0) = \mu_0(x)$. Using this notation, the ATT is given by

eqnarray[eqnarray omitted — 82 chars of source]

and the NN matching estimator is given by

eqnarray[eqnarray omitted — 253 chars of source]

We first show that $\hat \tau$ is an asymptotically unbiased estimator for the ATT when $N_1$ is fixed and $N_0 \rightarrow \infty$, and we derive its asymptotic distribution in this setting. We consider the following assumptions on how the distribution of $Y(0)|X = x$ changes with $x$.

assumption[Distribution of $Y(0)|X = x$] (a) $\mu_0(x)$ is continuous, and (b) for any $h(y)$ continuous and bounded, $\tilde h(x) = \mathbb{E}[h(Y(0))|X=x]$ is continuous and bounded.

Assumption (ref)(a) states that the conditional expectation of $Y(0)$ with respect to $X = x$ is continuous in $x$, which is standard in the matching literature. The intuition behind Assumption (ref)(b) is that the conditional distribution of $Y(0)$ given $X=x$ changes “smoothly” with $x$. This guarantees that $Y(0)|(X= x_n)$ converges in distribution to $Y(0)|(X= x)$ if $x_n \rightarrow x$, as we show in Appendix Lemma (ref).\footnote{We use this condition to apply the Portmanteau Lemma in the proof of Appendix Lemma (ref). Other equivalent conditions could be used. } In Appendix (ref), we show that this condition is satisfied if, for example, $Y(0)|(X=x) \sim N(\theta(x),\sigma(x))$, where $\theta(x)$ and $\sigma(x)$ are continuous functions of $x$.

For each $x \in \mathbb{X}_1$, let $\xi_x \sim (Y(1) - \mu_1(x)) | (X=x,W=1)$ and $\eta_x \sim (Y(0) - \mu_0(x)) | (X=x,W=0)$. Moreover, let $G(\kappa; x)$ be the CDF of $(\mu_1(x) - \mu_0(x) - \tau) +\xi_x - \frac{1}{M} \sum_{m=1}^M \eta_x^m$, where $\{\eta_x^m \}_{m=1}^M$ are iid copies of $\eta_x$, and $(\xi_x,\eta^1_x,\cdots,\eta^M_x)$ is mutually independent.

proposition(1) Under Assumptions (ref), (ref), (ref), and (ref)(a), $\mathbb{E}[\hat \tau ] \rightarrow \tau$ when $N_0 \rightarrow \infty$ and $N_1$ is fixed. (2) Under Assumptions (ref), (ref), (ref), and (ref)(b), \begin{eqnarray} \nonumber \hat \tau \buildrel d \over \rightarrow \tau + \frac{1}{N_1} \sum_{i \in \mathcal{I}_1} \kappa_i when $N_0 \rightarrow \infty$ and $N_1$ is fixed, \end{eqnarray} where the CDF of $\kappa_i$ is given by $\widetilde G(\kappa) = \int_{x \in \mathbb{X}_1} G(\kappa; x)f_1(x)dx$, and $\mathbb{E}[\kappa_i]=0$. Moreover, $\{\kappa_i\}_{i \in \mathcal{I}_1}$ is mutually independent.

Let $X^i_{(m)}$ be the covariate value of the $m$-closest match to observation $i \in \mathcal{I}_1$. The main intuition for the results in Proposition (ref) is that, for a fixed $X_i=\bar x$, $X^i_{(m)} \buildrel p \over \rightarrow \bar x$ when $N_0 \rightarrow \infty$, because, holding $M$ fixed, we will always be able to find $M$ observations in the control group that are arbitrarily close to $\bar x$. Independence of $\kappa_i$ follows from the fact that the probability of two treated observations sharing the same nearest neighbor converges to zero. See details in Appendix (ref).

Proposition (ref) shows that the expected value of the NN matching estimator converges to $\tau$. We also derive in Proposition (ref) the asymptotic distribution of $\hat \tau$, which has expected value equal to $\tau$. Therefore, the NN matching estimator is asymptotically unbiased whether we define asymptotic unbiasedness as $\mathbb{E}[\hat \tau] \rightarrow \tau$, or as $\hat \tau \buildrel d \over \rightarrow \tau + \widetilde \kappa$, with $\mathbb{E}[\widetilde \kappa]=0$.

remark\normalfont With $N_1$ fixed, the estimator is not consistent. This happens because, with $N_1$ fixed, we cannot apply a law of large numbers to the average of the error of the treated observations. For the same reason, the matching estimator will not generally be asymptotically normal. These conclusions are similar to the ones derived by CT for differences-in-differences estimators with few treated groups.
remark\normalfont Consider a bias-corrected estimator suggested by AI_2011, \begin{eqnarray} \hat \tau_{biasadj} = \frac{1}{N_1} \sum_{i \in \mathcal{I}_1} \left[ Y_i - \frac{1}{M} \sum_{m=1}^M \left( Y_j + \hat \mu_0(X_i) - \hat \mu_0(X^i_{(m)}) \right) \right], \end{eqnarray} where $\hat \mu_0(x)$ is an estimator for $\mu_0(x) $, and $X^i_{(m)}$ be the covariate value of the $m$-closest match to observation $i$. If the conditions on $\hat \mu_0(x)$ considered by AI_2011 are satisfied, then we can also guarantee that $\hat \tau_{biasadj} $ has the same asymptotic distribution as $\hat \tau$. The intuition is that $\hat \mu_0(X_i) - \hat \mu_0(X^i_{(m)})$ converges in probability to zero when $N_0 \rightarrow \infty$, because $X^i_{(m)} \buildrel p \over \rightarrow X_i$.
remark\normalfont We consider an asymptotic framework in which $M$ is held fixed, while $N_0 \rightarrow \infty$, which is similar to what AI call fixed-$M$ asymptotics in their setting. As argued by AI, the motivation for such fixed-$M$ asymptotics is to provide an approximation to the sampling distribution of matching estimators with a small number of matches. Matching estimators using few matches have been widely used in applied work (see AI). Moreover, Imbens_Rubin argue against using matching estimators with many matches, as this would tend to increase the bias of the resulting estimator, while the marginal gains in precision of increasing the number of matches are limited. In Section (ref) we consider the implication of our findings for other types of matching estimators.

Inference

The fact that the NN matching estimator is not generally asymptotically normal when $N_1$ is fixed and $N_0 \rightarrow \infty$ poses an important challenge when it comes to inference. In particular, inference based on the asymptotically normal distribution derived by AI, or on the bootstrap procedure suggested by Otsu, should not provide a good approximation in our setting, as the asymptotic theory behind these methods relies on both $N_1$ and $N_0$ going to infinity. We therefore consider alternative inference methods based on the theory of randomization tests under an approximate symmetry assumption, developed by Canay. We focus on a test based on sign changes, while in in Appendix (ref) we consider an alternative test based on permutations. We consider the problem of testing the null hypothesis $H_0: \tau = c$.

Without loss of generality, let $i=1,...,N_1$ be the treated observations, and consider a function of the data given by

eqnarray[eqnarray omitted — 87 chars of source]

where $\hat \tau_i^{N_0} = \left(Y_i - \frac{1}{M} \sum_{j \in \mathcal{J}_M(i)} Y_j \right)- c$. Each $\hat \tau_i^{N_0}$ depends on the $M$ nearest neighbors of observation $i$, so its distribution depends on $N_0$.

Following Canay, we consider a test statistic given by

eqnarray[eqnarray omitted — 130 chars of source]

where $\hat \tau^{N_0} = \frac{1}{N_1} \sum_{i \in \mathcal{I}_1} \hat \tau_i^{N_0}$. Note that $\hat \tau^{N_0} = \hat \tau$ if $c=0$.

We consider the group of transformations given by $\textbf{G} = \{ -1,1 \}^{N_1}$, where $gS_{N_0} = \left( g_1\hat \tau_1^{N_0},...,g_{N_1}\hat \tau_{N_1}^{N_0} \right)'$. Let $K = |\textbf{G}|$ and denote by

eqnarray[eqnarray omitted — 85 chars of source]

the ordered values of $\{ T(gS_{N_0}) : g \in \textbf{G} \}$. Let $k=\lceil K(1-\alpha) \rceil$, where $\alpha$ is the significance level of the test. Then the test is given by

eqnarray[eqnarray omitted — 179 chars of source]

In words, we calculate the test statistic $T(gS_{N_0} )$ for all possible $gS_{N_0} = \left( g_1\hat \tau_1^{N_0},...,g_{N_1}\hat \tau_{N_1}^{N_0} \right)'$, and then we compare the actual test statistic $T(S_{N_0} )$ with the distribution $\{ T(gS_{N_0} ) : g \in \textbf{G} \}$. We first show validity of such test when $N_0 \rightarrow \infty$ and $N_1$ is fixed under symmetry conditions on the distribution of potential outcomes and on the distribution of heterogeneous treatment effects.

assumption[Symmetry] (i) $Y(w) | (X=x,W=w)$ is symmetric around its mean for $w \in \{0,1\}$ and for all $x \in \mathbb{X}_1$, and (ii) the distribution of $\mu_1(X) - \mu_0(X)$ conditional on $W=1$ is symmetric around $ \tau$.

While this is a strong assumption, the condition that potential outcomes, conditional on $X$, are symmetric can be justified in settings in which observation $i$ is the average of a large number of individuals, by appealing to some central limit theorem.\footnote{Notice that this does not preclude dependence between individuals in observation $i$, insofar as it is still amenable to a central limit theorem.} This could be the case, for example, in our empirical application in which each observation represents average test scores of a large number of students per school, even if we do not observe student-level data. While we cannot test the plausibility of this assumption for $Y(1)$ (given that we have fixed $N_1$) and for the distribution of heterogeneous treatment effects, we can provide evidence on whether the distribution of $Y(0) | (X=x,W=0)$ is symmetric by fitting a model for $\mu_0(x)$ and checking whether the residuals are symmetric for the controls.

We show that the sign-changes test is asymptotically valid when $N_0 \rightarrow \infty$ under such symmetry assumptions, even when $N_1$ is fixed. Our MC simulations presented in Section (ref) suggest that relaxing Assumption (ref) does not generate large size distortions for this test, except in settings in which $N_1$ is very small, and the asymmetry in the potential outcomes or heterogeneous effects is very strong.

propositionSuppose Assumptions (ref), (ref), (ref), (ref)(b), and (ref) hold. Assume also that the distribution of $Y$ is continuous. If we consider the problem of testing $H_0: \tau =c$, then, for any $\alpha \in (0,1)$, $\mbox{limsup}_{N_0 \rightarrow \infty} \mathbb{E} \left[ \phi( S_{N_0}) \right] \leq \alpha$ when $N_1$ is fixed.

The main idea of the proof is to show that the limiting distribution of $S_{N_0}$, under the null, is invariant to the transformations in ${\textbf{G}}$. This is true because, asymptotically, $\hat \tau_i^{N_0}$ and $\hat \tau_j^{N_0}$ are independent for $i \neq j$, and, under the null, $\hat \tau_i^{N_0}$ converges in distribution to $\kappa_i$, which is symmetric around zero given Assumption (ref). Details in Appendix (ref).

While Proposition (ref) provides a test that is valid when $N_1$ is fixed, validity even when $N_1$ is fixed comes at a cost of relying on stronger assumptions than usually considered in the matching literature. We show that we can relax Assumption (ref) if $N_1$ increases, but at a slower rate relative to $N_0$. We consider the following assumptions, which are similar to the ones considered by AI for the setting in which $N_0$ grows at a faster rate than $N_1$.

assumption[Sampling rates] For some $r > \mbox{max}\{k/2,2\}, N_{1}^{r} / N_{0} \rightarrow \theta$ with $0<\theta<\infty$.
assumption[Distribution of potential outcomes] For $w=0,1,$ (i) $\mu(x, w)$ and $\sigma^{2}(x, w)$ are Lipschitz in $\mathbb{X}_0$, (ii) for some $\gamma>0$, $\mathbb{E}\left[ |\epsilon|^{4+\gamma} | W=w,X = x \right]$ exists and is bounded uniformly in $x,$ and (iii) $\sigma^{2}(x, w)$ is bounded away from zero.

Under these conditions, Corollary 1(ii) from AI implies that the NN matching estimator is consistent and asymptotically normal. We show that the sign-changes test is asymptotically valid in this setting in which $N_1$ increases (but at a lower rate than $N_0$) even when we relax Assumption (ref).

propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. If we consider the problem of testing $H_0: \tau =c$, then the sign-changes test is asymptotically valid when $N_1,N_0 \rightarrow \infty$.

Therefore, the sign-changes test is asymptotically valid under weaker conditions when $N_1$ also increases. The condition that $N_0$ grows at a faster rate than $N_1$ (Assumption (ref)) is important for two reasons. First, it guarantees that we can apply Corollary 1(ii) from AI, which implies that the bias of the NN matching estimator is asymptotically negligible. Second, it also guarantees that the probability that we have shared nearest neighbors converges to zero when $N_1,N_0 \rightarrow \infty$. See details of the proof in Appendix (ref).

remark\normalfont We can construct confidence intervals considering the set of values $c$ such that the null would not be rejected. Propositions (ref) and (ref) provide the assumptions we need for validity of such confidence intervals whether we consider a setting with $N_1$ fixed or $N_1 \rightarrow \infty$.
remark\normalfont In Propositions (ref) and (ref), this test is asymptotically valid because the probability that different treated observations share the same nearest neighbor goes to zero, when $N_0 \rightarrow \infty$. If there are shared nearest neighbors in finite samples, this may lead to over-rejection if we do not take that into account. Therefore, we suggest a finite sample adjustment, in which we restrict to sign changes such that $g_i=g_j$ if $i$ and $j$ share the same nearest neighbor. The probability that this modification is relevant converges to zero when $N_0 \rightarrow \infty$.\footnote{Another alternative would be to consider a matching estimator without replacement. However, this would generate lower quality matches, which implies more bias (AI). Moreover, matching without replacement has the disadvantage that the estimator is not invariant to different sorting of the data. }
remark\normalfont Canay consider a randomized version of the test to deal with cases such that $ T( S_{N_1}) = T^{( k)}( S_{N_1})$, while we consider a test that rejects if $ T ( S_{N_1}) > T^{( k)}( S_{N_1})$. Such randomization guarantees an asymptotic size of $\alpha$ even when $N_1$ is fixed.
remark\normalfont This test is also asymptotically valid for bias-corrected matching estimators, as defined in equation ((ref)). We define $\tilde \tau_{i}^{N_0} = Y_i - \frac{1}{M} \sum_{j \in \mathcal{J}_M(i)} \left( Y_j - \hat \mu_0(X_j) + \hat \mu_0(X_i) \right)$ in this case. We also consider the implications of Proposition (ref) for other types of matching estimators in Section (ref).

Monte Carlo Simulations

We present two sets of MC simulations. First, we present an empirical MC simulation, which provides a setting in which there is selection on observables with a structure based on a real application. Then we consider another set of MC simulations where the focus is to evaluate the relevance of Assumption (ref) for the sign-changes test when $N_1$ is small.

Empirical Monte Carlo simulations

We construct an empirical MC simulation in which treatment assignment and potential outcomes are based on the “Jovem de Futuro” program, which we present in more details as an empirical illustration in Section (ref).\footnote{More details on the “Jovem de Futuro” program and on the construction of this empirical MC study are also presented in Appendix (ref).} The main results of the empirical MC simulation are summarized in Table (ref). Panel A shows that, when we consider NN matching estimators with few nearest neighbors, the bias of the matching estimator is close to zero, regardless of the number of treated observations. This is true in our simulations even when the number of control observations is not large. Increasing the number of nearest neighbors used in the estimation implies that we need an increasing number of controls to keep our approximations reliable. We show in Appendix Table (ref) that increasing the dimensionality of the matching variables also implies that a larger number of controls is needed to keep our approximations reliable.

Panels B and C present rejection rates, respectively, for the asymptotic test based on AI and for the sign-changes test.\footnote{In Appendix Table (ref) we consider the test based on permutations presented in Appendix (ref) and the wild bootstrap test proposed by Otsu as alternative inference methods.} The asymptotic test generally presents over-rejection when $N_1$ is small, which is consistent with the fact that the theory behind this test relies on $N_1 \rightarrow \infty$. In contrast, the sign-changes test controls well for size even when $N_1$ is very small. An important caveat, however, is that the sign-changes test may be conservative in settings in which the number of sign-changes transformations is very small, which is a common feature in approximate randomization tests ART. The number of sign-changes transformations will be small when $N_1$ is very small, or when $N_0$ is small relative to $N_1 \times M$.\footnote{The number of sign-changes transformations can be small when $N_0$ is small relative to $N_1 \times M$ due to the finite-sample adjustment discussed in Remark (ref). } The sign-changes test presents non-trivial power, except for the cases in which it is very conservative (Appendix Table (ref)).

When $N_0$ is large, the sign-changes test presents non-trivial power in these simulations when $N_1=5$ because we consider a 10%-level test. However, if we considered a 5%-level test, then it would be very conservative and have a very low power in this case (Appendix Table (ref)). This happens because we would have very few sign-changes transformations to reject the null at a 5% significance level. An alternative in case we want to consider a 5%-level test with a very small number of treated observations is the approximate randomization test based on permutations, presented in Appendix (ref). Similarly to the sign-changes test, the test based on permutations (with the right choice of test statistic) is also valid under stronger assumptions when $N_1$ is fixed, and under weaker assumptions when $N_1$ increases (but at a lower rate than $N_0$). However, it relies on arguably stronger conditions than the sign-changes test when $N_1$ is fixed. We discuss that in more detail in Section (ref).

Monte Carlo simulations relaxing symmetry conditions

We consider now another set of MC simulations in which we vary de degree of symmetry of the potential outcomes and of the distribution of heterogeneous treatment effects. For all simulations, we set $(X | W=w) \sim N(0,1)$ for $w \in \{ 0,1\}$ and $Y(0) | (X= x,W=0) \sim N(0,1)$ for all $x \in \mathbb{R}$. Therefore, $\mu_0(x)=0$ for all $x \in \mathbb{R}$. Then we vary $\mu_1(x)$ and the distribution of $(Y(1) - \mu_1(x))|(X=x,W=1)$. For all settings, we consider $N_1 \in \{5,10,25,50 \}$ and $N_0 = 1000$.

We start considering a setting with $\mu_1(x) = x$ and $(Y(1) - \mu_1(x))|(X=x,W=1)\sim N(0,1)$, which implies that the symmetry conditions from Assumption (ref) hold. In this case, we have that $\tau = \mathbb{E}[\mu_1(X) - \mu_0(X) | W=1] = 0$, so the null hypothesis is true. We present rejection rates for 10% tests in Panel A of Table (ref). Consistent with the MC simulations from Section (ref), the asymptotic test based on AI over rejects when $N_1$ is small. In contrast, the sign-changes test does not over-reject irrespectively of $N_1$. This is expected, because Assumption (ref) is valid in this case, and we consider a setting in which the estimator is unbiased.

In Figure (ref), we contrast the power of these two tests when $M=4$, for different values of $N_1$. We modify the DGP so that $\mu_1(x) = \tau + x$, implying that $\mathbb{E}[\mu_1(X) - \mu_0(X)|W=1] = \tau$. We present size-adjusted power for the AI test using critical values that set rejection rates equal to 10% in the MC simulations when the null is true. Since the sign-changes test never over-rejects in this setting, rejection rates are not adjusted for this test. When $N_1=5$, the sign-changes test presents non-trivial power, but its power is lower than the size-adjusted power of the asymptotic test. We recall, however, that the asymptotic test presents large over-rejection in this scenario, so it is not a feasible alternative. When $N_1 = 10$, the loss in power of the sign-changes test is very small, while the asymptotic test still presents relevant over-rejection.

When we further increase $N_1$, the size distortion of the asymptotic test diminishes. In this case, the sign-changes test continues to control for size, but presents a small loss in power relative to the asymptotic test when $N_1$ increases. This happens because, in this case, $N_0$ becomes smaller relative to $N_1 \times M$. Consider the case in which $N_1=50$, $N_0=1000$, and $\tau = 0.5$. In this case, the size-adjusted power of the asymptotic test is $0.756$, while the power of the sign-changes test is $0.690$. If we increase $N_0$ to $2000$, then the gap in power between these tests goes down from 5.6pp to 3.7pp. In contrast, if we reduce $N_0$ to 500, then the gap in power increases to 11pp. Therefore, the sign-changes test does not present relevant losses in terms of power if $N_0$ is large, but may present some loss in power if $N_0$ is not much larger than $N_1 \times M$.

In panels B to E in Table (ref), we consider variations in the DGP in which the symmetry conditions from Assumption (ref) does not hold. We first set $\mu_1(x) = Q^{-1}(\Phi(x);p)-p)/\sqrt{2p}$, where $Q(x;p)$ is the CDF of a chi-squared distribution with $p \in \mathbb{N}$ degrees of freedom, and $\Phi(x)$ is the CDF of a standard normal. In this case, we have that $(\mu_1(X) - \mu_0(X)) | W=1$ has mean zero and variance one, but its distribution is asymmetric. The asymmetry is decreasing with $p$. We also consider settings in which $(Y(1) - \mu_1(x))|(X=x,W=1) \sim (\chi^2_p - p)/\sqrt{2p}$ instead of standard normal, which adds more asymmetry in the distribution of $\kappa_i$. Again, this distribution has mean zero and variance one, but it has an asymmetry that is decreasing in $p$. In Appendix Table (ref), we also consider cases in which $\mu_1(x)=0$ and we only vary the distribution of $(Y(1) - \mu_1(x))|(X=x,W=1)$. In these simulations, the sign-changes test continues to control for size, except when $N_1$ is small and the distribution of $\kappa_i$ is very asymmetric. Importantly, even in the scenarios in which the sign-changes test presents some over-rejection, its over-rejection is milder relative to the over-rejection of the asymptotic test. When $N_1$ increases, then both tests control for size, which is consistent with the fact that they are asymptotically valid when $N_1 \rightarrow \infty$, even when $\kappa_i$ is asymmetric.

Overall, based on these simulations, the sign-changes test presents important gains in terms of test size when $N_1$ is small, even when the distribution of $\kappa_i$ is asymmetric. Except when this distribution is extremely asymmetric, the sign-changes test does not present much over-rejection. Moreover, even when it presents some over-rejection in these simulations, the over-rejection is milder relative to the over-rejection of the asymptotic test. Finally, in settings in which the asymptotic test controls well for size, the cost in terms of power for the sign-changes test is low, as long as $N_0$ is sufficiently large relative to $N_1 \times M$. Therefore, this test provides an interesting alternative for settings in which $N_1$ is small, and also when $N_1$ is not very small, but $N_0 >> N_1$.

Comparing Alternative Inference Methods

The different test procedures we consider potentially present important trade-offs in terms of size distortion and power, depending on the number of treated and control observations. Moreover, the sign-changes test relies on different sets of assumptions depending on whether the empirical application is better approximated by a theory in which $N_1$ is fixed or in which $N_1$ diverges (but at a slower rate relative to $N_0$). In light of the theoretical properties derived in Section (ref), and of the evidence from the MC simulations presented in Section (ref), we provide guidance on how to consider the suitability of different inference methods in empirical applications.

If $N_1$ is large, then the asymptotic approximations considered by AI should be reliable. In this case, if we are in a setting in which $N_0$ is much larger than $N_1$, then the sign-changes test would be comparable to the asymptotic test. More specifically, both tests would be valid under the same assumptions regarding the distributions of potential outcomes, and they would have similar power. However, if $N_0$ is not very large relative to $M \times N_1$, then the sign-changes test may have lower power, and the asymptotic test should be preferable.

If $N_1$ is not very large, then the test based on AI presents relevant size distortions, and the sign-changes test becomes an interesting alternative. In this case, one should be aware that this test is valid under stronger assumptions if $N_1$ is very small, so these assumptions should be discussed by applied researchers. Since this test is asymptotically valid even when we relax such symmetry conditions when $N_1$ increases, we expect that distortions in case such assumptions are not valid to be relatively minor, except in cases in which $N_1$ is very small and errors or treatment effects are very asymmetric. This intuition is corroborated by the simulations presented in Section (ref). In those simulations, the sign-changes test only presents relevant size distortions when $N_1$ is very small and the degree of asymmetry in the distribution of $\kappa_i$ is large. Moreover, in such settings, the asymptotic test based on AI presents more severe size distortions than the sign-changes test.

Overall, if $N_1$ is not very large, then the sign-changes test presents relevant gains relative to the asymptotic test in terms of controlling for test size. Moreover, if $N_0$ is large relative to $N_1 \times M$, then the sign-changes test has a power comparable to the (size-adjusted) power of the asymptotic test. {The only exception in which the sign-changes test would have a lower power than the asymptotic test even when $N_0$ is large is when $N_1$ is very small (for example, when $N_1=5$). However, those are exactly the cases in which we should expect the size distortions of the asymptotic test to be more severe. }

Finally, it is worth noting that the sign-changes test only presents non-trivial power in settings in which $N_1$ is very small (say, $N_1 = 5$), when we consider 10%-level tests. A feasible alternative when $N_1$ is even smaller than 5, or when we want to consider a 5%-level test, is the test based on permutations described in Appendix (ref). However, one should be aware that, with $N_1$ fixed, such test would rely on homoskedasticity and treatment effects homogeneity assumptions, which are arguably stronger than Assumption (ref). Similarly to the sign-changes test, the test based on permutations, with the right choice of test statistic, is also valid under weaker assumptions when $N_1$ increases (but at a lower rate than $N_0$).

Empirical Illustration

As an empirical illustration of the NN matching estimator in a setting with small $N_1$ relative to $N_0$, we analyze the “Jovem de Futuro” program. This is a program that has been running in Brazil since 2008, aimed at improving the quality of education in public schools by improving management practices and allocating grants to treated schools. In 2010, this program was implemented in a randomized control trial with 15 treated schools in Rio de Janeiro and 39 treated schools in Sao Paulo, with the same number of control schools in each state. In Appendix (ref) we present more details on this empirical application. We rely on this randomized control trial to validate the use of NN matching estimators in a setting in which there are few treated and many control units.\footnote{Influential papers that evaluate the use of non-experimental methods in empirical applications where a randomized control trial is available include 10.2307/2677743, Heckman13416, 10.2307/2999630, 10.2307/2971733, ASMITH2005305, Lalonde, DW1999, and DW2002.} We take advantage of the fact that there were about 1,000 other public schools in Rio de Janeiro and more than 3,000 other public schools in Sao Paulo that did not participate in the experiment. More specifically, we consider a NN matching estimator using the experimental control schools as treated observations, and schools that did not participate in the experiment as control observations. These experimental control schools were selected following the same process used for the selection of treated schools. However, since these schools did not actually receive the treatment in the analyzed period, we should not expect to find significant effects in this case if the matching estimator is valid 10.2307/2677743. Therefore, this provides an interesting setting to evaluate the validity of matching estimators with few treated and many control observations.

Table (ref) shows estimated effects from 2010 to 2012. We use test scores from 2007 to 2009 as matching variables. In addition to the point estimates, p-values are calculated using the asymptotic distribution derived by AI, and from the sign-changes test. Interestingly, estimates for Rio de Janeiro (columns 1 to 4) generally have lower p-values using the test based on AI, relative to the alternative inference procedure. In particular, a test based on the asymptotic distribution would reject the null at 10% in two cases, while the sign-changes test would fail to reject the null. This is consistent with our simulations from Section (ref), where we show that the asymptotic test based on AI may lead to over-rejection when $N_1$ is small. The difference in p-values across different methods is less pronounced when we consider estimates for Sao Paulo, which is consistent with having a larger number of “treated” schools in Sao Paulo.

Conclusion

We consider the asymptotic properties of matching estimators when the number of control observations is large, but the number of treated observations is fixed. In this setting, the NN matching estimator is asymptotically unbiased for the ATT under standard assumptions used in the literature on estimation of treatment effects under selection on unobservables. Moreover, we provide tests, based on the theory of randomization under approximate symmetry, that are asymptotically valid when the number of treated observations is fixed and the number of control observations goes to infinity. While we need to rely on relatively strong assumptions so that these tests are valid even when $N_1$ is fixed, we show that these tests are also valid under weaker assumptions when $N_1$ increases, but at a slower rate relative to $N_0$. We analyze in details the advantages and disadvantages of these inference methods, and provide guidance on which methods should be used in specific applications.

We conjecture that the asymptotic unbiasedness and the asymptotic validity of the randomization inference test based on sign changes when $N_1$ is fixed remain valid if we consider other types of matching estimators. Intuitively, the main requirement should be that, when constructing the counter-factual for a treated unit $j \in \mathcal{I}_1$, the estimator would rely on a weighted average of the control observations such that, as $N_0 \rightarrow \infty$, an increasing proportion of the weights would be allocated to control units $i$ with $X_i$ close to $X_j$. This would be true if we use, for example, a kernel method for the weights under suitable conditions on the smoothing parameter 10.2307/2971733. In this case, the bias of the treatment effect estimator for each treated observation $j \in \mathcal{I}_1$ would go to zero as $N_0 \rightarrow \infty$. Moreover, the correlation between the treatment effect estimator for different treated observations would also go to zero in this case, providing asymptotic validity for the inference method based on sign changes. However, an advantage of considering the NN matching estimator in this setting is that it would be more straightforward to implement the adjustment proposed in Remark (ref) to avoid over-rejection in finite samples for the inference method based on sign changes. Moreover, it would also be possible to consider the randomization test based on permutations, presented in Appendix (ref), when we rely on NN matching estimators.

Our results are also relevant for synthetic control (SC) applications. Following Doudchenko, the SC and the matching estimators are nested in a framework in which the estimated counterfactual outcome for the treated observation is a linear combination of the outcomes for the controls. In their framework, consider an estimator in which the weights given to control observations with large discrepancies in pre-treatment outcomes relative to the treated units go to zero. In this case, following the same arguments as above, the estimator would be asymptotically unbiased if treatment assignment is “as good as random,” conditional on this set of pre-treatment outcomes.\footnote{See Abadie2010, FB, FP_SC, and ferman for a discussion on the validity of the synthetic control estimator under a different set of assumptions.} This is exactly the case for the penalized SC estimator for disaggregated data proposed by lhour. Under these conditions, the randomization inference test we propose based on sign changes remains asymptotically valid when the number of control units goes to infinity. This provides an interesting alternative for inference, when there are multiple treated units and a large number of control units, that does not rely on exchangeability nor homoskedasticity assumptions.\footnote{See Firpo2018SyntheticCM, FP_tests and Hahn for a discussion on the placebo test proposed by Abadie2010. Victor propose a permutation test based on the timing of the intervention. This test, however, would require a very large number of periods. Instead, our test may be an alternative when the number of periods is not large, but the number of control units is large. } The only caveat is that a very large number of control observations is needed when the number of pre-treatment periods is large, so that approximations remain reliable.

\singlespace

table[table omitted — 2,080 chars of source]
table[table omitted — 3,119 chars of source]
figure[figure omitted — 1,153 chars of source]
table[table omitted — 1,813 chars of source]