EconBase
← Back to paper

Specification Tests for the Propensity Score

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

98,409 characters · 15 sections · 117 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Specification Tests for the Propensity Score

abstract\onehalfspacing{ This paper proposes new nonparametric diagnostic tools to assess the asymptotic validity of different treatment effects estimators that rely on the correct specification of the propensity score. We derive a particular restriction relating the propensity score distribution of treated and control groups, and develop specification tests based upon it. The resulting tests do not suffer from the \textquotedblleft curse of dimensionality\textquotedblright\ when the vector of covariates is high-dimensional, are fully data-driven, do not require tuning parameters such as bandwidths, and are able to detect a broad class of local alternatives converging to the null at the parametric rate $n^{-1/2}$, with $n$ the sample size. We show that the use of an orthogonal projection on the tangent space of nuisance parameters facilitates the simulation of critical values by means of a multiplier bootstrap procedure, and can lead to power gains. The finite sample performance of the tests is examined by means of a Monte Carlo experiment and an empirical application. Open-source software is available for implementing the proposed tests. \newline JEL: C12, C31, C35, C52. \newline Keywords: Empirical Processes; Integrated Moments; Multiplier Bootstrap; Projection; Treatment Effects. }

Introduction

The propensity score, which is defined as the conditional probability of receiving treatment given covariates, is one of the most widely used tools for causal inference. Part of its popularity can be credited to the seminal result of Rosenbaum1983: if the treatment assignment is independent of the potential outcomes conditional on a vector of covariates, then one can obtain unbiased and consistent estimators of different treatment effect measures by adjusting for the propensity score alone, greatly reducing the dimensionality of the underlying problem. Several methods that exploit this important insight are now an essential part of the applied researcher's toolkit. Examples include matching, see e.g. Rosenbaum1985, Heckman1997 and Abadie2016; inverse probability weighting (IPW), see e.g. Rosenbaum1987a, Hirano2003 and Donald2013; regression methods, see e.g. Hahn1998 and Firpo2007; and many others. For literature reviews, see Heckman2007 and Imbens2009.

Despite their popularity, a main concern of these methods is that the propensity score is usually unknown, and therefore has to be estimated. Given the high dimensionality of available covariates, researchers are usually coerced to adopt a parametric model for the propensity score since nonparametric estimation methods suffer from the \textquotedblleft curse of dimensionality\textquotedblright , implying that the resulting treatment effect estimators can have considerably poor properties, even for large sample sizes. Such a common practice raises the important issue of model misspecification. Indeed, as shown by Frolich2004, Millimet2009 , Huber2013 and Busso2014, propensity score misspecifications can lead to misleading treatment effect estimates.

In this paper we propose new specification tests for parametric propensity score models. Our proposal builds on the common practice of comparing the density of the propensity score between treated and control groups to determine the covariate overlap region, see e.g. Heckman1998. However, instead of comparing conditional densities, we focus on comparing conditional cumulative distribution functions (CDFs). In particular, we derive a restriction between the propensity score CDFs among treated and control groups that gives information on overlapping\footnote{ We thank an anonymous referee for making this suggestion.}, show that such a restriction is equivalent to a particular infinite number of unconditional moment conditions, and develop tests based upon it.

In contrast to existing proposals, our tests are fully data-driven, do not require user-chosen tuning parameters such as bandwidths, and are able to detect a broad class of local alternatives converging to the null at the parametric rate $n^{-1/2}$, with $n$ the sample size. Furthermore, our tests do not suffer from the \textquotedblleft curse of dimensionality\textquotedblright\ when the vector of covariates is of high dimensionality, and have greater power than competing tests for many alternatives. Of course, such power gains do not come without a cost: there exist some classes of alternative hypotheses against which our tests have trivial power. Nonetheless, we believe that such a compromise is reasonable since, as pointed out by Janssen2000 and Escanciano2009a, achieving reasonable power over all possible directions seems hopeless.

The proposal closest to ours is Shaikh2009. Despite using a similar characterization of the null hypothesis as Shaikh2009, our proposal greatly differs from theirs. Whereas Shaikh2009 adopts the local smoothing approach, see e.g. Hardle1993, Zheng1996, Fan1996 and Li1998, we adopt the integrated conditional moment (ICM) approach, see e.g. Bierens1982, Bierens1990, Bierens1997 , Stute1997 and Escanciano2006b. As a consequence, our approach inherits some advantages of ICM when compared to Shaikh2009. First, our tests do not require delicate bandwidth choice, unlike Shaikh2009's test whose performance can be sensitive to it. Second, in contrast with Shaikh2009, our approach has power against local alternatives converging to the null at the parametric rate.

Another popular procedure to assess misspecification of the propensity score model is to use \textquotedblleft balancing\textquotedblright\ tests. Initially proposed by Rosenbaum1985, these tests consist of assessing if each covariate is independent of the treatment assignment, conditional on the propensity score. This is often implemented examining whether moments (usually just the mean) of the observable characteristics between the two \textquotedblleft matched\textquotedblright\ or \textquotedblleft weighted\textquotedblright\ groups are the same; see e.g. Dehejia2002 and Smith2005. One should bear in mind that because \textquotedblleft balancing\textquotedblright\ tests are usually based on a finite number of orthogonality conditions, there are uncountably many directions of misspecification that cannot be detected with these tests. Furthermore, as shown by Lee2013a, balancing tests may have size distortions due to the \textquotedblleft multiple testing problem\textquotedblright , the failure to account for the estimation effect of the propensity score, and poor covariate overlap. Such drawbacks put at stake the reliability of many of these procedures. Our proposal, on the other hand, does not suffer from these.

Our paper also contributes to the literature on ICM tests. What appears distinctive to our approach is that $(i)$ we exploit the dimension-reduction coming from our derived restriction between propensity score CDFs, and $ \left( ii\right) $ we acknowledge our lack of knowledge of the \textquotedblleft true\textquotedblright\ correct specification of the propensity score, by means of an orthogonal projection onto the tangent space of nuisance parameters. The result of $(i)$ and $(ii)$ is a test with improved power properties, and with a simple bootstrap implementation. The power improvement due to the dimension reduction has been noticed by Stute2002, Escanciano2006b and Shaikh2009, whereas the power improvement due to the use of orthogonal projections has been noticed in different contexts, see e.g. Neyman1959, and more recently, Bickel2006 and Escanciano2014a. To the best of our knowledge, our proposal is the first to incorporate both procedures.

As mentioned above, our paper is related to the relatively scarce literature on projection-based specification tests, see e.g. Escanciano2009simple and Escanciano2014a for two notable exceptions. Escanciano2009simple proposes a simple bootstrap testing procedure for conditional moment restrictions that acknowledges specifically the fact that the nuisance parameters are unknown and introduces the projection methodology, while Escanciano2014a propose a projection-based testing procedure for linear quantile regression using a related projection weight function. Our proposal builds on these papers, with the important difference that our test statistics also exploit the dimension-reduction property of the propensity score.

The rest of the paper is organized as follows. In Section (ref) we present the testing framework and derive the restriction upon which our tests are based. The asymptotic properties of our tests are established in Section (ref). We next examine the finite sample properties of our tests by means of a Monte Carlo study in Section (ref). We provide an empirical illustration of our procedures in Section (ref). Section (ref) concludes. Mathematical proofs are gathered in an appendix at the end of the article.

Finally, all proposed tests discussed in this article can be implemented via open-source R package pstest, which is freely available from GitHub (\url{https://github.com/pedrohcgs/pstest}).

Testing Framework

Background

Let $D$ be a binary random variable that indicates participation in the program, i.e. $D=1$ if the individual participates in the treatment and $D=0$ otherwise. Define $Y\left( 1\right) $ and $Y\left( 0\right) $ as the potential outcomes under treatment and control, respectively. The realized outcome of interest is $Y=DY\left( 1\right) +\left( 1-D\right) Y\left( 0\right) $, and $X$\ is an observable $d\times 1$ vector of pre-treatment covariates.\ Denote the support of $X$ by $\mathcal{X\subseteq }\mathbb{R}^{d}$ and the propensity score $p\left( x\right) =\mathbb{P} \left( D=1|X=x\right) $. We have a random sample $\left\{ \left( Y_{i},D_{i},X_{i}^{\prime }\right) ^{\prime }\right\} _{i=1}^{n}$ of size $ n\geq 1$ from $\left( Y,D,X^{\prime }\right) ^{\prime }$. Throughout the rest of this article, all random variables are defined on a common probability space $\left( \Omega ,\mathcal{A},\mathbb{P}\right) .$

The main goal in causal inference is to assess the effect of a treatment $D$ on the outcome of interest $Y$. The most popular parameters of interest include the average treatment effect, $ATE=\mathbb{E}\left[ Y\left( 1\right) -Y\left( 0\right) \right] $, and the average treatment effect on the treated, $ATT=\mathbb{E}\left[ Y\left( 1\right) -Y\left( 0\right) |D=1\right] $. Note that such parameters of interest depend on potential outcomes $ Y\left( 1\right) $ and $Y\left( 0\right) $ which cannot be jointly observed for the same individual, precluding estimating $ATE$ and $ATT$ using their sample analogues. One of the most popular identification strategies in policy evaluation that resolves such difficulty is to assume that selection into treatment is solely based on observable characteristics, the so-called unconfoundedness setup, see e.g. Rosenbaum1983. Formally, the unconfoundedness setup requires the following assumptions:

assumption$\left( Y\left( 1\right) ,Y\left( 0\right) \right) \protect\mathpalette{\protect\independentT}{\perp} D|X$.
assumption$\forall x\in \mathcal{X}$, $0<p\left( x\right) <1.$

As shown by Rosenbaum1987a, under Assumptions (ref)-(ref), $ATE$ and $ATT$ are identified by

equation*[equation* omitted — 306 chars of source]

respectively. This result motivates the two-step procedure in which one first estimates the propensity score, computes its estimated values $\hat{p} \left( X_i\right) $, and then uses the analogy principle to estimate $ATE$ and $ATT$, that is,

eqnarray*[eqnarray* omitted — 407 chars of source]

Alternatively, one could estimate $ATE$ and $ATT$ using propensity score matching, see e.g. Rosenbaum1983, Heckman1998a and Abadie2016.

In order to ensure that such estimators are well-defined and stable, it is important to assess the overlap between the distribution of the propensity score among treatment and control groups, i.e. to check whether the propensity score is bounded away from zero and one, and if the support of the propensity score in both groups are nearly the same, see e.g. Heckman1998, Smith2005, Crump2009 and Khan2010. Following Heckman1998, it is now routine to compare kernel density estimates of the propensity score among treated and control samples to determine the common support region. In cases where there is strong overlap one proceeds as described above, otherwise, one usually considers trimmed samples, see e.g. Crump2009 and Sasaki2018.

Although kernel density estimators are popular, they involve choosing tuning parameters such as bandwidths and often suffer from boundary bias. Of course such inconveniences can be easily avoided if one focuses on CDFs instead of densities. In the following we show that propensity score overlap implies a particular set of restrictions between the CDFs of treated and control groups, and that these restrictions can form the basis for testing the correct specification of propensity score models.

Assume that the propensity score $p\left( X\right) $ has a density with respect to a dominating measure, and that the density is bounded away from zero and infinity uniformly over its support. The following lemma builds on Shaikh2009 and formalizes the above discussion.

lemmaLet $\alpha =\mathbb{P}\left( D=0\right) /\mathbb{P}\left( D=1\right) $ and assume that $0<\mathbb{P}\left( D=1\right) <1$. If $ 0<p\left( X\right) <1~a.s.,$ then \begin{equation} \mathbb{E}\left[ 1\left\{ p\left( X\right) \leq u\right\} |D=1\right] =\alpha \mathbb{E}\left[ \frac{p\left( X\right) }{1-p\left( X\right) } 1\left\{ p\left( X\right) \leq u\right\} |D=0\right], \forall u\in \left[ 0,1 \right] . \end{equation} Furthermore, ((ref)) holds if and only if \begin{equation} \mathbb{E}\left[ \left( D-p\left( X\right) \right) 1\left\{ p\left( X\right) \leq u\right\} \right] =0, \forall u\in \left[ 0,1\right] . \end{equation}

Lemma (ref) implies that, when the propensity score is correctly specified, one can expect that the sample analogue of ((ref)) should hold. Thus, ((ref)) provides a graphical diagnostic tool for propensity score misspecification; see Lemma 3.2 of Sloczynski2018 for a result related to ((ref)). Perhaps more importantly, note that ((ref)) provides an infinite number of simple unconditional moment restrictions that can be used to formally test whether a parametric model for the propensity score is correctly specified or not.

Motivated from Lemma (ref), we seek to test whether a parametric putative model for $p\left( x\right) $ is correctly specified based on

equation[equation omitted — 187 chars of source]

where $\Theta \subset \mathbb{R}^{k}$, $\Pi=\left[ 0,1\right]$ is the unit interval, and $q\left( X,\theta \right) :$ $\mathcal{X\times }\Theta \mapsto \left[ 0,1\right] $ is a family of parametric functions known up to the finite dimensional parameter $\theta $. Common specifications for $q\left( X,\theta \right) $ in empirical applications are the Probit, $\Phi \left( X^{\prime }\theta \right) ,$ and the Logit, $\Lambda \left( X^{\prime }\theta \right) $, where $\Phi \left( \cdot \right) $ and $\Lambda \left( \cdot \right) $ are the normal and logistic link functions, respectively.

Note that ((ref))\ can be equivalently written as

equation[equation omitted — 176 chars of source]

see e.g. Stute1997\footnote{ Alternative representations of $H_{0}$ are also possible, see e.g. Bierens1997, Escanciano2006a and a previous version of this article. }. Thus, in order to assess $H_{0}$, one can either use the infinite number of unconditional moment restrictions in ((ref)), or the conditional moment restriction in ((ref)). In this article we use ((ref)) whereas Shaikh2009 exploits ((ref)). More precisely, Shaikh2009 consider a test statistic based on

equation[equation omitted — 348 chars of source]

where $\varepsilon _{i}\left( \hat{\theta}_{n}\right) =D_{i}-q\left( X_{i}, \hat{\theta}_{n}\right) $, $\hat{\theta}_{n}$ is a $\sqrt{n}-$consistent estimator of $\theta _{0}$ under $H_{0}$, $h_{n}$ is a positive scalar bandwidth parameter converging to zero at a suitable rate as $n\rightarrow \infty $, and $K\left( \cdot \right) $ is a kernel function. Note that, in addition to the estimation of $\theta _{0}$ under $H_{0}$, Shaikh2009 's procedure requires local smoothing of the data, implying that its finite sample properties rely on the adequate choice of the smoothing parameter $ h_{n}$, a task that is far from trivial in testing problems. Since our approach is based on ((ref)) and only involves unconditional expectations, our testing procedure is free of tuning parameters such as bandwidth sequence $h_{n}$.

Tests based on a continuum of unconditional moment restrictions such as ((ref)) fall into the ICM approach, see Gonzalez-Manteiga2013 for a review. Nonetheless, our tests have two main differences with respect to the standard ICM tests. First, ((ref)) depends on $X$ only through the propensity score model under $H_{0}$, a one-dimensional (though unknown) function. As a consequence, the ICM in ((ref)) is insensitive to the dimension $d$ of the explanatory variables $X$, avoiding the so-called \textquotedblleft curse of dimensionality\textquotedblright . Second, in contrast to the standard ICM tests, we explicitly acknowledge that $\theta _{0}$ is a nuisance parameter in testing ((ref)) by proposing to use orthogonal projections on the tangent space of nuisance parameters. As discussed in the Introduction, the use of orthogonal projections leads to important advantages. In the next subsection we describe how we construct such projection-based tests, paying particular attention to the role played by the orthogonal projection; see also Escanciano2009simple and Escanciano2014a for related results in different contexts.

RemAlthough the identification of treatment effect parameters such as the $ATE$ relies on both Assumptions (ref) and (ref) , the results in Lemma (ref) (and the null hypothesis ((ref))) do not involve outcome data and are therefore well motivated even when Assumption (ref) may not hold. This situation may arise in decomposition exercises, see e.g. Fortin2011 for a review of decomposition methods in economics. With respect to Assumption ((ref)), we note that it allows propensity scores to be arbitrarily close to zero and one and hence it is not very restrictive. In fact, given that the unconditional moment condition in ((ref) ) does not involve random denominators, weak covariate overlap does not play a major role in our testing procedure. However, weak covariate overlap may lead to irregular treatment effect estimators, see e.g. Khan2010.

Projection-based specification tests

Recall that $\varepsilon _{i}\left( \theta \right) =D_{i}-q\left( X_{i},\theta \right) .$ For all $u\in \Pi $, define

equation[equation omitted — 212 chars of source]

where $g(x,\theta )=\partial q(x,\theta )/\partial \theta$ is the score function of $q(x,\theta )$,

equation*[equation* omitted — 116 chars of source]

and

equation*[equation* omitted — 119 chars of source]

Given a random sample $\left\{ \left( D_{i},X_{i}^{\prime }\right) ^{\prime }\right\} _{i=1}^{n}$, our test statistics are based on continuous functionals of the projection-based empirical process $\hat{R}_{n}^{p}(u)$,

equation[equation omitted — 194 chars of source]

where $\hat{\theta}_{n}$ is a $\sqrt{n}-$consistent estimator for $\theta _{0}$ under $H_{0}$. Two popular examples of such functionals are the Cram \'{e}r-von Mises-type and Kolmogorov-Smirnov-type functionals,

align[align omitted — 299 chars of source]

respectively, where $F_{n}(u)=n^{-1}\sum_{i=1}^{n}1\left( q\left( X_{i},\hat{ \theta}_{n}\right) \leq u\right) $ is the empirical distribution function (EDF) of $q\left( X_{i},\hat{\theta}_{n}\right) $, $1\leq i\leq n.$

At this point, one may wonder why our test statistics are based on the empirical process $\hat{R}_{n}^{p}(u)$ ((ref)) instead of the usual sample analogue of ((ref)),

equation[equation omitted — 167 chars of source]

i.e. the \textquotedblleft unprojected\textquotedblright\ analogue of $\hat{R }_{n}^{p}(u)$. To answer such a query, note that under $H_{0}$ and some weak regularity conditions given in Section (ref), the unprojected process $\hat{R}_{n}(u)$ can be decomposed as

multline[multline omitted — 311 chars of source]

uniformly in $u\in \Pi $, see Lemma (ref) in the Appendix. The asymptotic representation in ((ref)) implies that the effect of replacing $\theta _{0}$ by $\hat{\theta}_{n}$ is non-negligible, and therefore the asymptotic null distributions of tests based on ((ref)) are sensitive to the estimator $\hat{\theta}_{n}$ being used. As a consequence, for a given parametric specification $p\left( x\right) =q(X,\theta _{0})$, the asymptotic null distributions of tests based on ((ref)) will depend on whether one estimates $\theta _{0}$ using maximum likelihood (ML), nonlinear least squares (NLS), or generalized method of moments (GMM), even though the underlying specification for the propensity score is the same across these estimation methods.

The projection-based process $\hat{R}_{n}^{p}(u)$, on the other hand, avoids such drawback since

equation[equation omitted — 130 chars of source]

almost everywhere in $u\in \Pi $, where

equation[equation omitted — 206 chars of source]

with

equation*[equation* omitted — 104 chars of source]

and

equation*[equation* omitted — 107 chars of source]

The intuition behind ((ref)) is simple. First, note that $\Delta ^{-1}\left( \theta \right) G\left( u,\theta \right) $ is the vector of linear projection coefficients of regressing $1\left\{ q(X,\theta )\leq u\right\} $ on $g(X,\theta )$. Thus, it follows that $g(X,\theta )^{\prime }\Delta ^{-1}\left( \theta \right) G\left( u,\theta \right) $ is the best linear predictor of $1\left\{ q(X,\theta )\leq u\right\} $ given $g(X,\theta )$, and that ((ref)) is nothing more than the associated projection error, which is by definition orthogonal to $g(X,\theta )$. As a consequence of ((ref)), it follows that under some weak regularity conditions, uniformly in $u\in \Pi $,

equation*[equation* omitted — 72 chars of source]

where

equation[equation omitted — 162 chars of source]

see Theorem (ref) in Section (ref). Thus, $\hat{R}_{n}^{p}(u)$ is (asymptotically) invariant to the choice of estimator $\hat{\theta}_{n}$. Furthermore, as we discuss in Section (ref), the above asymptotic representation of $\hat{R}_{n}^{p}(u)$ in terms of $R_{n0}^{p}(u)$ allows for a multiplier-type bootstrap procedure that greatly simplifies the computation of asymptotically valid critical values.

In summary, by focusing on the projection-based process $\hat{R}_{n}^{p}(u)$ instead of the more traditional process $\hat{R}_{n}(u),$ our proposed test statistics are $\left( a\right) $ invariant to the choice of estimator $\hat{ \theta}_{n}$, and $\left( b\right) $ allow for simplified bootstrap implementation. In addition, tests based on $\hat{R}_{n}^{p}(u)$ acknowledge that deviations in the direction of the score function $g(x,\theta)$ cannot be distinguished from deviations within the parametric model, and therefore do not \textquotedblleft waste\textquotedblright\ power in such directions. As a result, tests based on $\hat{R}_{n}^{p}(u$) can have higher power when compared to tests based on $\hat{R}_{n}(u)$, though in general none of them is strictly better than the other uniformly over the space of alternatives. We defer the discussion of these power properties to Section (ref).

Asymptotic theory

In this section, we establish the asymptotic behavior of the projection-based empirical process $\hat{R}_{n}^{p}(u)$ under the null hypothesis $H_{0}$, under the fixed alternative hypothesis $H_{1}$, which is the negation of ((ref)), and under a sequence of local alternatives $ H_{1n}$ that converges to $H_{0}$ at the parametric rate $n^{-1/2}$, $n$ being the sample size. We also characterize classes of alternative hypotheses against which our tests have no power, and argue that such classes are rather exceptional. Finally, we show that critical values can be computed with the assistance of a multiplier-type bootstrap that is easy to implement.

Asymptotic null distribution

The asymptotic null distributions of our tests are the limiting distributions of continuous functionals of $\hat{R}_{n}^{p}(u)$ under $H_{0}$ . To derive the asymptotic results, we adopt the following notation. For a generic set $\mathcal{G}$, let $l^{\infty }\left( \mathcal{G}\right) $ be the Banach space of all uniformly bounded real functions on $\mathcal{G}$, equipped with the uniform metric $\left\Vert f\right\Vert _{\mathcal{G} }\equiv \sup_{z\in \mathcal{G}}\left\vert f\left( z\right) \right\vert $. We study the weak convergence of $\hat{R}_{n}^{p}(u)$ and its related processes as elements of $l^{\infty }\left( \Pi \right) $, where $\Pi \equiv \left[ 0,1 \right] $. Let “$\Rightarrow $” denote weak convergence on $\left( l^{\infty }\left( \Pi \right) ,\mathcal{B}_{\infty }\right) $ in the sense of J. Hoffmann-J$\phi $rgensen, where $\mathcal{B}_{\infty }$ denotes the corresponding Borel $\sigma $-algebra - see e.g. Definition 1.3.3 in VanderVaart1996.

We assume the following regularity conditions. Let $\Theta _{0}$ be an arbitrarily small neighborhood around $\theta _{0}$ such that $\Theta _{0}\subset \Theta $. For any $d_1\times d_2$ matrix $A=(a_{ij})$, let $ ||A|| $ denote its Euclidean norm, i.e. $||A||=[\text{tr}(AA^{\prime})] ^{1/2}$.

assumption$(i)$ The parameter space $\Theta $ is a compact subset of $ \mathbb{R}^{k};$ $\left( ii\right) $ the true parameter $\theta _{0}$ belongs to the interior of $\Theta $; and $\left( iii\right) $ $\left\Vert \hat{\theta}_{n}-\theta _{0}\right\Vert =O_{p}(n^{-1/2}).$
assumptionThe parametric propensity score function $q(x,\theta )$ is twice continuously differentiable in $\Theta _{0}$ for each $x\in\mathcal{X}$ , with its first derivative $g(x,\theta )=\partial q(x,\theta )/\partial \theta =(g_{1}(x,\theta ),\ldots ,g_{k}(x,\theta ))^{\prime }$ satisfying $ \mathbb{E}[\sup_{\theta \in \Theta _{0}}||g(X,\theta )||]<\infty $ and its second derivative satisfying $\mathbb{E}[\sup_{\theta \in \Theta _{0}}||\partial g(X,\theta )/\partial \theta ||]<\infty $. Furthermore, the matrix $\Delta (\theta )\equiv \mathbb{E}[g(X,\theta )g^{\prime }(X,\theta )] $ is nonsingular in $\Theta _{0}$.
assumptionThe function $F_{\theta }(u)=\mathbb{P}(q(X,\theta )\leq u)$ satisfies $\sup_{u\in \Pi }|F_{\theta _{1}}(u)-F_{\theta _{2}}(u)|\leq C||\theta _{1}-\theta_{2}||$, where $C$ is a bounded positive number, not depending on $\theta_1$ and $\theta_2$.

Assumptions (ref)-(ref) are weaker than related conditions in the literature. For instance, Assumption (ref) only requires $\sqrt{n} \left( \hat{\theta}_{n}-\theta _{0}\right) =O_{p}\left( 1\right) ,$ but does not require $\sqrt{n}\left( \hat{\theta}_{n}-\theta _{0}\right) $ to admit an asymptotically linear representation. Assumption (ref) is a condition concerning the degree of smoothness of the propensity score $ q(x,\theta )$, and is satisfied for standard parametric models such as the Probit and the Logit specifications. It also only requires finite first moment of $g(X,\theta )$, instead of more than four moments as in Shaikh2009. Assumption (ref) simply imposes a Lipschitz type continuity condition on the CDF of the parametric propensity score.

RemAssumption (ref) is used to prove that the class of functions $\mathcal{ F}=\{x\mapsto 1\left\{ q(x,\theta )\leq u\right\} :u\in \Pi ,\,\theta \in \Theta \}$ is Donsker, see Lemma (ref) in the Appendix. Such assumption is similar to condition (5.8) in Lee2011e. Alternatively, if $\mathcal{F}_{\theta }=\{x\mapsto q(x,\theta ):\theta \in \Theta \}$ is a VC class of functions, the aforementioned Donsker result also follows even without Assumption (ref), see e.g. Example 2.1 in VanderVaart2007 .

Next, we derive the asymptotic behavior of the projection-based empirical process $\hat{R}_{n}^{p}(u)$ under $H_{0}$. We do this in two steps. First, we show that, under $H_{0}$, $\hat{R}_{n}^{p}(u)$ is asymptotically equivalent, with respect to the supremum norm on $\Pi $, to the process $ R_{n0}^{p}(u)$ given in ((ref)). From this result it follows that the weak convergence under $H_{0}$ of the process $\hat{R}_{n}^{p}(u)$ can be conveniently established from that of $R_{n0}^{p}(u)$. More importantly, the limiting null behavior of $\hat{R}_{n}^{p}(u)$ does not depend on $\hat{ \theta}_{n}$ nor how $\hat\theta_n$ is obtained.

theoremLet Assumptions (ref)-(ref) hold. Then, under $H_{0}$, we have that \begin{equation*} \sup_{u\in \Pi }\left\vert \hat{R}_{n}^{p}(u)-R_{n0}^{p}(u)\right\vert =o_{p}(1), \end{equation*} and \begin{equation*} \hat{R}_{n}^{p}(u)\Rightarrow R_{\infty }^{p}, \end{equation*} where $R_{\infty }^{p}$ denotes a Gaussian process with mean zero and covariance structure given by \begin{equation} K^{p}(u_{1},u_{2})=\mathbb{E}\left[q(X,\theta _{0})\left(1-q(X,\theta _{0})\right)\mathcal{P}1\left\{q(X,\theta _{0})\leq u_{1}\right\}\mathcal{P} 1\left\{q(X,\theta _{0})\leq u_{2}\right\}\right]. \end{equation}

Theorem (ref) and the continuous mapping theorem (CMT), see e.g. Theorem 1.3.6 in VanderVaart1996, yield the asymptotic null distributions of continuous functionals of $\hat{R}_{n}^{p}(u),$ including the test statistics $CvM_{n}$ and $KS_{n}$ given in ((ref)) and ((ref) ), respectively.

corollaryUnder the assumptions of Theorem (ref) and $H_{0}$, for any continuous functional $\Gamma (\cdot )$ from $l^{\infty }\left( \Pi\right) $ to $\mathbb{R}$, we have \begin{equation*} \Gamma (\hat{R}_{n}^{p})\xrightarrow{d}\Gamma (R_{\infty }^{p}). \end{equation*} Furthermore, \begin{equation*} CvM_{n}\xrightarrow{d}CvM_{\infty }:=\int_{\Pi }\left\vert R_{\infty }^{p}(u)\right\vert ^{2}\,dF_{\theta _{0}}(u), \end{equation*} where $F_{\theta _{0}}(u)=\mathbb{P}\left( q\left( X,\theta _{0}\right) \leq u\right) $ denotes the cumulative distribution function of $q\left( X,\theta _{0}\right)$, and \begin{equation*} KS_{n}\xrightarrow{d}KS_{\infty }:=\sup_{u\in \Pi }\left\vert R_{\infty }^{p}(u)\right\vert . \end{equation*}

Note that the integrating measure in $CvM_{n}$ is a random measure, but Corollary (ref) shows that the asymptotic distribution is not affected by this fact. Further details can be found in the Appendix A.

Asymptotic power

Now, we investigate the power properties of tests based on continuous functionals $\Gamma (\hat{R}_{n}^{p})$, like $CvM_{n}$ and $KS_{n}$ in ((ref)) and ((ref)), respectively. We consider fixed alternatives, and a sequence of local alternatives $H_{1n}$ that converges to $H_{0}$ at the parametric rate $n^{-1/2}$.

Power against fixed alternatives

Next theorem analyzes the asymptotic properties of our tests under fixed alternatives of the type

equation[equation omitted — 175 chars of source]

where $\Pi =[0,1]$ is the unit interval. Note that $H_{1}$ is simply the negation of $H_{0}$ in ((ref)).

theoremSuppose Assumptions (ref)-(ref) hold. Then, under the fixed alternative hypothesis $H_{1}$ in ((ref)), we have that \begin{equation*} \sup_{u\in \Pi }\left\vert \frac{1}{\sqrt{n}}\hat{R}_{n}^{p}(u)-\mathbb{E} \left[ \left( p\left( X\right) -q\left( X,\theta _{0}\right) \right) \mathcal{P}1\left\{ q\left( X,\theta _{0}\right) \leq u\right\} \right] \right\vert =o_{p}\left( 1\right) . \end{equation*}

From Theorem (ref), we see that test statistics of the form of $\Gamma ( \hat{R}_{n}^{p})$ are not consistent against all fixed alternative hypotheses in ((ref)), but only those not collinear to the score function $g(X,\theta_0)$. To see this, note that

multline*[multline* omitted — 497 chars of source]

is equal to zero under ((ref)) if $p\left( X\right) -q\left( X,\theta _{0}\right)$ and $g(X,\theta _{0})$ are collinear almost surely. We do not see this as a limitation. First, when one estimates $\theta _{0}$ using the NLS method, the population first order condition for $\theta _{0}$ sets $ \mathbb{E}\left[ \left( D-q\left( X,\theta _{0}\right) \right) g^{\prime }(X,\theta _{0})\right] =0,$ implying that, for some $u\in \Pi ,$

equation*[equation* omitted — 315 chars of source]

As a consequence, our projection-based tests would be consistent against all alternative hypotheses of the type of ((ref)), avoiding the aforementioned problem.

Second, and perhaps more importantly, even when one does not use NLS to estimate $\theta _{0},$ we argue that the lack of power against alternatives collinear to the score function $g(X,\theta _{0})$ is not a main concern. As shown by Escanciano2009a, every test based on ICM approach has trivial local power against these alternatives, and as a consequence, the global power of all ICM tests in the direction of the score function will also be low, see e.g. Strasser1990. In fact, instead of considering this as a limitation one may consider this property as a feature of our tests: by acknowledging that such alternatives cannot be powerfully detected, our projection-based test statistics do not waste power in such directions, and therefore may have higher power against other, perhaps more important, alternatives; see Section (ref) for an illustration.

The above discussion raises the question: are there other classes of fixed alternative hypotheses that our specification tests are not able to detect? As first pointed out by Shaikh2009, the answer is yes. Because our test statistics depend on $X$ only through the propensity score $q(X,\theta _{0})$, our tests will have trivial power against the class of misspecified propensity scores, where

equation*[equation* omitted — 130 chars of source]

in a set of positive probability, but

equation*[equation* omitted — 116 chars of source]

where $\theta _{0}$ now is the probability limit of $\hat{\theta}_{n}$. In other words, tests based on ((ref)) (or equivalently on ((ref))), will have trivial power against alternatives such that $p\left( X\right) $ is different from $q\left( X,\theta _{0}\right) $ but

equation[equation omitted — 144 chars of source]

The leading case of such a class of alternatives is when the propensity score is correctly specified for a subvector of $X$, but not for the entire vector $X$, see e.g. Shaikh2009. Given the nonlinear nature of $ p\left( X\right) $, one may consider such a class of alternatives rather exceptional. However, they can still arise in practice under some particular circumstances as we will illustrate below.

Consider the case where $X=\left( X_{1},X_{2}\right) $ and that both covariates are relevant for the propensity score, i.e., $p\left( X\right) =p\left( X_{1},X_{2}\right) $. Suppose that a researcher considers a Probit model for the $p\left( \cdot \right) $ but only included $X_{1}$ as a covariate, i.e., the researcher assume the model $q\left( X,\theta _{0}\right) =\Phi \left( \theta _{00}+\theta _{01}X_{1}\right) $, with $\Phi \left( \cdot \right) $ the normal link function. Of course, $p\left( X\right) \not=q\left( X,\theta _{0}\right) $ since $q\left( X,\theta _{0}\right) $ only includes a subset of relevant covariates. Thus, in light of ((ref)), our tests will have no power if there exists $\theta _{0}=\left( \theta _{00},\theta _{01}\right)^{\prime }\in \Theta \subset \mathbb{R}^2$ such that $\mathbb{E}\left[ p\left( X_{1},X_{2}\right) |X_{1} \right] =\Phi \left( \theta _{00}+\theta _{01}X_{1}\right) $ $a.s.$. The existence of such $\theta _{0}$ depends on the underlying conditional distribution of $X_{2}$ given $X_{1}$, and also on the form of true unknown propensity score $p\left( X_{1},X_{2}\right) $. For instance, assume that the true propensity score is a Probit, e.g., $p\left( X_{1},X_{2}\right) =\Phi \left( X_{1}+X_{2}\right) $, and that $X_{2}$ given $X_{1}$ follows a normal distribution with conditional mean $\mu \left( X_{1}\right) $ and conditional variance $\sigma ^{2}\left( X_{1}\right) $. In this case, we can show that

equation*[equation* omitted — 197 chars of source]

Thus, if $\mu \left( X_{1}\right) =a+bX_{1}$ and $\sigma^2 \left( X_{1}\right) =c$ for some constants $a,b$ and $c$ with $b\neq -1$ and $c>0$, we have that

equation*[equation* omitted — 221 chars of source]

Thus, in this particular case such $\theta _{0}$ does exist with $ \theta_0\equiv\left(a/\sqrt{1+c},(1+b)/\sqrt{1+c}\right)^{\prime }$ and our tests would have trivial power. On the other hand, if $\mu \left( X_{1}\right) $ is nonlinear in $X_{1}$ and $\sigma ^{2}\left( X_{1}\right) $ is a nontrivial function of $X_{1}$, no such $\theta _{0}$ exists and therefore ((ref)) is ruled out. Thus, it is clear that the conditional distribution of $X_{2}$ given $X_{1}$ plays an important role in the \textquotedblleft empirical relevance\textquotedblright\ of alternative hypotheses like ((ref)).

The role played by the functional form of the true propensity score in the \textquotedblleft empirical relevance\textquotedblright\ of ((ref)) can be illustrated by assuming that the true propensity score is a Logit instead of a Probit, e.g., $p\left( X_{1},X_{2}\right) =\Lambda \left( X_{1}+X_{2}\right) $, with $\Lambda(\cdot)$ the logistic link function. In this case, because of the nonlinear nature of $\Lambda $, there exists no $ \theta _{0}\in \Theta $ such that $\mathbb{E}\left[ \Lambda \left( X_{1}+X_{2}\right) |X_{1}\right] =\Phi \left( \theta _{01}+\theta _{01}X_{1}\right) $ $a.s.$, ruling out ((ref))\footnote{ In practice, however, the power against this particular alternative can be low since the Logit and Probit specifications are relatively \textquotedblleft close\textquotedblright\ to each other.}.

In summary, the above discussion shows that our projection-based tests are consistent against a broad range of alternatives, though not all. This is the main drawback of our proposal when compared to the standard ICM specification tests. However, in our simulations that follow we show that, for the alternatives considered, our projection-based tests are the best or comparable to the best tests in terms of power in finite samples. Thus, from a practical point of view, the benefits of using our procedure can outweigh the costs in many relevant situations.

Power against local alternatives

Next, we study the performance of our projection-based tests under a sequence of local alternative hypotheses converging to the null at the parametric rate $n^{-1/2}$ given by

equation[equation omitted — 147 chars of source]

for some $\theta _{0}\in \Theta $, where $r\left(q(X,\theta _{0})\right)$ represents directions of departure from $H_0$, and $n^{-1/2}$ indicates the rate of convergence of $H_{1n}$ to $H_0$. The function $r:\left[ 0,1\right] \rightarrow \mathbb{R}$ is required to satisfy the following assumption.

assumptionThe function $r(q)$ is continuous in $q$ and satisfies $\mathbb{ E}|r(q(X,\theta _{0}))|<\infty $.
theoremSuppose Assumptions (ref) -(ref) hold. Then, under the local alternatives $H_{1n}$ given by ((ref)), we have \begin{equation*} \hat{R}_{n}^{p}(u)\Rightarrow R_{\infty }^{p}+\Delta _{r}, \end{equation*} where $R_{\infty }^{p}$ is the same Gaussian process as defined in Theorem (ref), and $\Delta _{r}$ is a deterministic shift function given by \begin{equation*} \Delta _{r}(u)\equiv \mathbb{E}\left[r\left( q\left( X,\theta _{0}\right) \right) \mathcal{P}1\left\{ q(X,\theta _{0})\leq u\right\} \right]. \end{equation*}

Note that, in general, the deterministic shift function $\Delta _{r}\left( u\right) \not=0$ for at least some $u\in \Pi $, implying that tests based on continuous even functionals of $\hat{R}_{n}^{p}(\cdot )$ will have non-trivial power against local alternatives of the form in ((ref)). A situation in which our tests will have trivial local power against such alternatives is when directions $r(q\left(x,\theta _{0}\right) )$ are a linear combination of score function $g(x,\theta _{0})$, i.e. $r(q(x,\theta _{0}))=\beta g(x,\theta _{0})$ for some $\beta $. In such a case, the limiting distribution of $\hat R_n^p(u)$ under $H_0$ and $H_{1n}$ is the same so that $H_{1n}$ cannot be detected. On the other hand, note that tests based on the local smoothing approach such as Shaikh2009 are not able to detect alternatives of the form ((ref)).

Computation of critical values

From the above theorems, we see that the asymptotic distribution of continuous functionals $\Gamma \left( \hat{R}_{n}^{p}\right) $ depend on the underlying data generating process and of course on $\Gamma \left( \cdot \right) $ itself. Furthermore, the complicated covariance structure of $ K^{p}(\cdot ,\cdot )$ given in ((ref)) does not allow for a simple representation of $R_{\infty }^{p}$ in terms of a well-known distribution-free Gaussian process for which critical values are readily available. To overcome this problem, we propose to compute critical values with the assistance of a multiplier bootstrap. The proposed procedure has good theoretical and empirical properties, is computationally easy to implement, and does not require computing new parameter estimates at each bootstrap replication.

More precisely, in order to estimate the critical values, we propose to approximate the asymptotic behavior of $\hat{R}_{n}^{p}(u)$ by that of

equation[equation omitted — 210 chars of source]

where $\{V_{i}\}_{i=1}^{n}$ is a sequence of $i.i.d.$ random variables with zero mean, unit variance and bounded support, independent of the original sample $\{(D_{i},X_{i}^{\prime })^{\prime }\}_{i=1}^{n}$. A popular example involves $i.i.d.$ Bernoulli variates $\left\{ V_{i}\right\} $ with $\mathbb{P }\left( V=1-\kappa \right) =\kappa /\sqrt{5}$ and $\mathbb{P}\left( V=\kappa \right) =1-\kappa /\sqrt{5}$, where $\kappa =\left( \sqrt{5}+1\right) /2,$ as suggested by Mammen1993.

With $\hat{R}_{n}^{p\ast }(u)$ at hands, the bootstrapped version of our test statistics $\Gamma \left( \hat{R}_{n}^{p}\right) $ is simply given by $ \Gamma \left( \hat{R}_{n}^{p,\ast }\right) $. For instance, the bootstrapped versions of $CvM_{n}$ and $KS_{n}$ in ((ref)) and ((ref)), respectively, are given by

align*[align* omitted — 234 chars of source]

The asymptotic critical values are then estimated by

equation*[equation* omitted — 253 chars of source]

where $\mathbb{P}_{n}^{^{\ast }}$ means bootstrap probability, i.e. conditional on the original sample $\{(D_{i},X_{i}^{\prime })^{\prime }\}_{i=1}^{n}$. In practice, $c_{n,\alpha }^{^{\Gamma ~\ast }}$ is approximated as accurately as desired by $\left( \Gamma \left( \hat{R} _{n}^{p,\ast }\right) \right) _{B(1-\alpha )}$, the $B\left( 1-\alpha \right)-$th order statistic from $B$ replicates $\left\{ \left(\Gamma \left( \hat{R}_{n}^{p,\ast }\right)\right)_l \right\} _{l=1}^{B}$ of $\Gamma \left( \hat{R}_{n}^{p,\ast }\right) $.

The next theorem establishes the asymptotic validity of the multiplier bootstrap procedure proposed above.

theoremAssume Assumptions (ref)-(ref). Then, \begin{equation*} \hat{R}_{n}^{p\ast }\underset{\ast }{\Rightarrow }R_{\infty }^{p}\quad a.s., \end{equation*} where $R_{\infty }^{p}$ is the Gaussian process defined in Theorem 1, and $ \underset{\ast }{\Rightarrow }$ denotes the weak convergence under the bootstrap law, i.e. conditional on the original sample $\{(D_{i},X_{i}^{ \prime })^{\prime }\}_{i=1}^{n}$. Additionally, for any continuous functional $\Gamma (\cdot )$ from $l^\infty(\Pi)$ to $\mathbb{R}$, we have $ \Gamma \left( \hat{R}_{n}^{p,\ast }\right) \underset{\ast }{\overset{d}{ \rightarrow }}\Gamma \left( R_{\infty }^{p}\right) $ a.s. under the bootstrap law.

Monte Carlo simulation study

In this section, we conduct a series of Monte Carlo experiments in order to study the finite sample properties of our proposed projection-based tests. In particular, we compare our Cram\'{e}r-von Mises and Kolmogorov-Smirnov tests $CvM_{n}$ and $KS_{n}$ given in ((ref)) and ((ref)) to $(i)$ the Shaikh2009's test,

equation*[equation* omitted — 160 chars of source]

where $\hat{V}_{n}\left( h_{n}\right) $ is given in ((ref)) and

equation*[equation* omitted — 350 chars of source]

the analogues of $CvM_{n}$ and $KS_{n}$ based on either $\left( ii\right) $ the unprojected process $\hat{R}_{n}(u)$ given in ((ref))\footnote{ For conciseness, we do not discuss the asymptotic properties of tests based on the unprojected empirical process $\hat{R}_{n}(u)$ in the main text. However, most of the asymptotic properties follow from arguments analogous to those we used to study $\hat{R}_{n}^{p}(u)$.}, or on $\left( iii\right) $ the traditional empirical process

equation[equation omitted — 162 chars of source]

We also compare our proposal with $\left( iv\right) $ balancing tests based on the normalized IPW estimators

equation[equation omitted — 178 chars of source]

where $w_{1,i}=D_{i}/q\left( X_{i},\hat{\theta}_{n}\right) $, $ w_{0,i}=\left( 1-D_{i}\right) /\left( 1-q\left( X_{i},\hat{\theta} _{n}\right) \right) $, $\bar{w}_{d,n}$ is the sample mean of $w_{d,i},$ $ d=\left\{ 0,1\right\} $, and $X^{j}$ is the $j$-th element of $d$ -dimensional vector $X=\left( X_{1},\dots ,X_{d}\right) ^{\prime }$. We consider two different test statistics for the balancing tests: the Wald test, and the maximum of the $d$ marginal two-sided $t$-tests.

Critical values (CVs) for the $CvM_{n}$ and $KS_{n}$ tests are obtained using the multiplier bootstrap procedure described in Section (ref), whereas for the ICM tests in $\left( ii\right) $ and $\left( iii\right) $ we use the bootstrap procedure described in Stute1998b. For $T_{n}\left( h_{n}\right)$ test, we use one-sided CVs from the standard normal distribution. CVs for the Wald test are from the chi-squared distribution with $d$ degrees of freedom. For the $t$-test, we consider Bonferroni corrected CVs based on the standard normal distribution. Note that the Bonferroni correction is necessary to address the multiple testing problem.

We consider sample sizes $n$ equal to $100$, $200$, $400$, $600$, $800$ and $ 1,000$. For each design, we consider $1,000$ Monte Carlo experiments. The $ \left\{ V_{i}\right\} _{i=}^{n}$ used in the bootstrap implementations are independently generated as $V$ with $\mathbb{P}\left( V=1-\kappa \right) =\kappa /\sqrt{5}$ and $\mathbb{P}\left( V=\kappa \right) =1-\kappa /\sqrt{5} $, where $\kappa =\left( \sqrt{5}+1\right) /2$, as proposed by Mammen1993. The bootstrapped critical values are approximated using $B=999$ bootstrap replications. To compute Shaikh2009's test, we use the standard normal kernel

equation*[equation* omitted — 101 chars of source]

Following Shaikh2009, the bandwidth sequence $h_{n}$ is chosen to be equal to $cn^{-1/8}$ for $c$ equal to 0.01, 0.05, 0.10 and 0.15. We choose different $c$'s to assess how sensitive Shaikh2009's test may be with respect to the bandwidth $h_{n}$.

Simulation 1

We first consider the following data generating processes (DGPs):

align*[align* omitted — 578 chars of source]

For each of these five DGPs, $D=1\left\{ D^{\ast }>0\right\} ,\varepsilon \protect\mathpalette{\protect\independentT}{\perp} \left( X_{1},X_{2}\right) $, where $X_{1}=Z_{1}$, $X_{2}=\left( Z_{1}+Z_{2}\right) /\sqrt{2}$, and $Z_{1}$, $Z_{2},$ and $\varepsilon $ are independent standard normal random variables. All the DGPs considered have the propensity score bounded away from zero and one. Finally, for each of these DGPs we consider the potential outcomes

equation*[equation* omitted — 143 chars of source]

where $m_{1}\left( X\right) =1+X_{1}+X_{2}$, $u\left( 1\right) $ and $ u\left( 0\right) $ are independent normal random variables with mean zero and variance $0.1$. The observed outcome is $Y=DY\left( 1\right) +\left( 1-D\right) Y\left( 0\right) $, and the true $ATE$ is 1. Although these outcome equations are not necessary to assess the size and power properties of the tests, they can be used to assess the utility of our proposed tests to distinguish between \textquotedblleft good\textquotedblright\ and \textquotedblleft bad\textquotedblright\ estimates of the $ATE$.

Let $X=\left( 1,X_{1},X_{2}\right) ^{\prime }$. For $DGP1-DGP5$, the $H_{0}$ considered is

equation*[equation* omitted — 230 chars of source]

where $\Phi (\cdot )$ is the cumulative distribution function (CDF) of the standard normal distribution. We estimate $\theta _{0}$ using the Probit ML, i.e.

equation*[equation* omitted — 228 chars of source]

Clearly, $DGP1$ falls under $H_{0}$, whereas $DGP2-DGP5$ fall under $H_{1}$. Note that $D$ follows a heteroskedastic Probit model in $DGP4$ and $DGP5$.

The simulation results are presented in Table (ref). We report empirical rejection frequencies at the $5\%$ significance level. Results for 10% and 1% significance levels are similar and are available upon request. We also report the bias of the normalized IPW estimator

equation[equation omitted — 153 chars of source]

and the average length and coverage of its estimated 95% confidence interval based on its asymptotically normal approximation (assuming that the propensity score model is correctly specified).

We first analyze the size of our test. From the results of $DGP1$, we find that the actual finite sample size of both $KS_{n}$ and $CvM_{n}$ tests is close to their nominal size, even when the sample size is as small as 100. The same holds for the other ICM-type tests. On the other hand, we find that Shaikh2009's test is, in general, conservative, and sensitive to the choice of bandwidth. For instance, when $c=0.15$, the empirical size is close to zero even with $n=1,000$. On the other hand, with $c=0.01$, the empirical size of Shaikh2009's test is closer to the nominal value. In terms of traditional balancing tests, we find that Wald test and Bonferroni-corrected $t$-test do not control size. Such a drawback is due to the random denominator being relatively close to zero: when we trim observations with estimated propensity score outside the $[0.05,0.95]$ range, we find that classical balancing tests can control size. We report these results in Table (ref) in the Appendix B. Finally, note that when the propensity score is correctly specified, the bias of the $ \widehat{ATE_n}$ estimator in ((ref)) is small, the length of the 95% confidence interval reduces as sample size increases, but the coverage probability is smaller than its nominal value even when $n=1,000$. However, as we show in Table (ref) in the Appendix B, such undercoverage disappears when we trim observations with extreme estimated propensity scores.

\afterpage{

landscape\begin{table}[htbp] \caption{Monte Carlo results under designs $DGP1$-$DGP5$} \begin{adjustbox}{ max width=1\linewidth, max totalheight=1\textheight, keepaspectratio} \begin{threeparttable} \begin{tabular}{ccccccccccccccccc} \hline \toprule DGP & $n$ & \multicolumn{1}{c}{$CvM_{n}$} & \multicolumn{1}{c}{$KS_{n}$} & \multicolumn{1}{c}{$T_{n}(0.01)$} & \multicolumn{1}{c}{$T_{n}(0.05)$} & \multicolumn{1}{c}{$T_{n}(0.10)$} & \multicolumn{1}{c}{$T_{n}(0.15)$} & \multicolumn{1}{l}{Max-$t$} & \multicolumn{1}{l}{$Wald$} & \multicolumn{1}{c}{$CvM_{n}^{unp}$} & \multicolumn{1}{c}{$KS_{n}^{unp}$} & \multicolumn{1}{c}{$CvM_{n}^{trad}$} & \multicolumn{1}{c}{$KS_{n}^{trad}$} & \multicolumn{1}{c}{Bias} & \multicolumn{1}{c}{CI length} & \multicolumn{1}{c}{Coverage} \\ \hline \midrule 1 & 100 & 5.00 & 5.60 & 4.90 & 2.30 & 0.80 & 0.20 & 8.90 & 6.90 & 5.40 & 5.00 & 5.40 & 6.10 & 0.11 & 1.36 & 85.60 \\ 1 & 200 & 4.40 & 4.60 & 4.90 & 2.30 & 0.60 & 0.20 & 8.80 & 8.70 & 4.00 & 4.50 & 4.40 & 3.90 & 0.02 & 1.04 & 90.00 \\ 1 & 400 & 5.10 & 5.80 & 4.10 & 2.40 & 1.10 & 0.40 & 12.10 & 12.40 & 4.70 & 5.30 & 4.70 & 4.80 & 0.01 & 0.76 & 90.80 \\ 1 & 600 & 5.20 & 5.70 & 4.30 & 1.80 & 0.70 & 0.40 & 13.20 & 14.50 & 5.00 & 5.30 & 4.60 & 5.50 & 0.01 & 0.64 & 87.20 \\ 1 & 800 & 5.30 & 4.30 & 4.50 & 1.50 & 0.50 & 0.40 & 13.80 & 16.30 & 4.90 & 4.50 & 6.20 & 4.70 & 0.01 & 0.56 & 89.50 \\ 1 & 1000 & 5.00 & 5.90 & 4.50 & 2.90 & 1.20 & 0.80 & 13.50 & 15.10 & 5.10 & 4.90 & 5.70 & 5.00 & 0.01 & 0.52 & 90.70 \\ \hline 2 & 100 & 28.90 & 26.40 & 6.70 & 7.60 & 7.70 & 6.20 & 14.00 & 26.80 & 17.90 & 18.70 & 16.10 & 16.80 & -1.76 & 6.48 & 89.00 \\ 2 & 200 & 61.50 & 55.30 & 7.30 & 17.30 & 20.80 & 18.70 & 13.80 & 27.00 & 41.60 & 42.60 & 35.50 & 30.50 & -2.75 & 6.39 & 78.70 \\ 2 & 400 & 90.60 & 84.30 & 18.50 & 44.60 & 50.40 & 50.40 & 41.60 & 45.00 & 79.80 & 74.80 & 71.50 & 60.60 & -3.50 & 5.97 & 29.30 \\ 2 & 600 & 98.90 & 96.30 & 33.60 & 66.60 & 74.00 & 75.80 & 75.50 & 75.50 & 94.90 & 93.40 & 89.80 & 82.80 & -3.95 & 5.84 & 9.80 \\ 2 & 800 & 99.70 & 99.10 & 50.40 & 83.50 & 89.20 & 89.70 & 88.40 & 89.50 & 99.10 & 98.00 & 97.70 & 93.80 & -4.14 & 5.78 & 3.60 \\ 2 & 1000 & 100.00 & 100.00 & 67.40 & 93.90 & 96.10 & 96.80 & 95.80 & 96.00 & 99.90 & 99.90 & 99.40 & 98.40 & -4.30 & 5.44 & 0.80 \\ \hline 3 & 100 & 32.80 & 26.10 & 6.20 & 5.30 & 1.50 & 0.40 & 0.10 & 0.10 & 33.60 & 25.00 & 15.90 & 18.10 & 0.01 & 0.87 & 95.90 \\ 3 & 200 & 59.10 & 49.90 & 18.70 & 16.50 & 6.40 & 1.10 & 0.00 & 0.00 & 59.00 & 49.60 & 40.80 & 38.80 & -0.01 & 0.57 & 94.70 \\ 3 & 400 & 69.20 & 64.10 & 38.40 & 28.60 & 9.20 & 2.40 & 0.00 & 0.00 & 69.70 & 63.70 & 83.00 & 80.30 & 0.00 & 0.39 & 94.70 \\ 3 & 600 & 79.70 & 74.60 & 56.60 & 40.90 & 14.70 & 3.80 & 0.00 & 0.00 & 80.10 & 75.30 & 98.70 & 97.10 & 0.00 & 0.32 & 95.70 \\ 3 & 800 & 81.00 & 77.70 & 63.40 & 42.30 & 15.20 & 4.10 & 0.00 & 0.00 & 81.30 & 77.80 & 99.70 & 99.70 & 0.00 & 0.27 & 96.20 \\ 3 & 1000 & 83.30 & 81.30 & 71.00 & 48.90 & 17.70 & 3.90 & 0.10 & 0.00 & 82.90 & 81.00 & 100.00 & 100.00 & 0.00 & 0.24 & 95.80 \\ \hline 4 & 100 & 15.20 & 12.20 & 4.40 & 1.50 & 1.10 & 0.50 & 0.30 & 0.30 & 15.60 & 12.00 & 20.00 & 11.40 & 0.17 & 0.90 & 90.70 \\ 4 & 200 & 35.00 & 27.00 & 5.10 & 7.30 & 6.50 & 3.50 & 1.10 & 0.60 & 37.50 & 26.80 & 42.00 & 25.10 & 0.16 & 0.60 & 80.60 \\ 4 & 400 & 64.70 & 53.80 & 8.10 & 17.20 & 21.10 & 17.00 & 4.10 & 2.20 & 70.30 & 51.20 & 74.60 & 46.30 & 0.16 & 0.42 & 69.20 \\ 4 & 600 & 85.30 & 75.40 & 13.40 & 33.00 & 39.00 & 35.70 & 8.10 & 5.20 & 88.30 & 72.50 & 90.50 & 69.90 & 0.16 & 0.34 & 55.50 \\ 4 & 800 & 92.50 & 86.70 & 19.00 & 50.00 & 60.40 & 59.40 & 9.50 & 7.40 & 94.80 & 85.00 & 96.70 & 83.70 & 0.16 & 0.29 & 41.90 \\ 4 & 1000 & 96.50 & 93.80 & 29.80 & 63.20 & 74.00 & 75.40 & 12.80 & 10.00 & 97.60 & 92.00 & 99.20 & 90.40 & 0.16 & 0.26 & 32.50 \\ \hline 5 & 100 & 11.00 & 7.80 & 5.00 & 2.20 & 1.30 & 0.60 & 22.40 & 22.40 & 9.20 & 6.90 & 9.20 & 5.70 & 0.04 & 2.70 & 68.30 \\ 5 & 200 & 16.40 & 13.70 & 4.90 & 2.70 & 2.30 & 1.60 & 29.10 & 32.40 & 13.80 & 9.70 & 13.00 & 7.80 & -0.31 & 2.74 & 68.00 \\ 5 & 400 & 26.70 & 23.00 & 5.90 & 5.70 & 4.70 & 3.90 & 24.10 & 29.60 & 23.50 & 15.80 & 22.70 & 11.10 & -0.72 & 3.04 & 72.60 \\ 5 & 600 & 35.80 & 32.80 & 5.30 & 8.70 & 9.90 & 8.40 & 18.50 & 24.50 & 35.20 & 21.70 & 32.40 & 15.50 & -0.93 & 3.26 & 77.40 \\ 5 & 800 & 45.00 & 39.30 & 5.70 & 11.80 & 14.30 & 12.90 & 17.50 & 21.80 & 44.00 & 28.60 & 38.80 & 19.10 & -1.02 & 3.23 & 77.80 \\ 5 & 1000 & 56.00 & 49.20 & 7.80 & 17.00 & 20.70 & 19.30 & 16.60 & 21.70 & 57.20 & 38.30 & 52.40 & 27.00 & -1.06 & 3.21 & 76.30 \\ \hline \bottomrule \end{tabular} \begin{tablenotes}[para,flushleft] { Note: Simulations based on 1,000 Monte Carlo experiments. “$CvM_{n}$” and “$KS_{n}$” stand for our proposed Cram\'{e}r-von Mises and Kolmogorov-Smirnov tests. “$T_{n}(c)$” stands for Shaikh2009's test, with bandwidth $h_{n}=cn^{-1/8}$. “Max-$t$” and “$Wald$” stand for, respectively, the Bonferroni-corrected Max-$t$-test, and Wald balancing test based on $\widehat{ATE}_{n}\left( X^{j}\right)$ defined in ((ref)). “$CvM_{n}^{unp}$” and “$KS_{n}^{unp}$” are the Cram\'{e}r-von Mises and Kolmogorov-Smirnov tests based on the unprojected empirical process $\hat{R}_{n}(u)$ defined in ((ref)), whereas “$CvM_{n}^{trad}$” and “$KS_{n}^{trad}$” are defined analogously, but based on $\hat{R}_{n}^{trad}\left( x\right)$ as defined in ((ref)). Finally, “Bias”, “CI length', and “Coverage” stands for the average simulated bias, estimated 95% confidence interval length, and 95% coverage probability for the $ATE$ estimator $\widehat{ATE_{n}}$ as defined in ((ref)). All entries are proportions of rejections at 5% level, in percentage points, except “Bias” ,“CI length”, and “Coverage” (measure in percentage points), which are as described above. See the main text for further details.} \end{tablenotes} \end{threeparttable} \end{adjustbox} \end{table}

}

Note that when the propensity score is misspecified in $DGP2-DGP5$, the $ATE$ estimator ((ref)) can be severely biased, and its confidence interval can be \textquotedblleft too small\textquotedblright\ when the propensity score is misspecified, leading to severe undercoverage\footnote{ In $DGP3$, the bias, confidence interval length, and coverage of the $ATE$ estimator are good. However, it is important to have in mind that we consider only one particular DGP for the $ATE$, and that such \textquotedblleft robust\textquotedblright\ results may not translate to other DGPs.}. Thus, detecting propensity score misspecifications can prevent misleading inference about $ATE$. Our proposed $KS_{n}$ and $CvM_{n}$ tests perform admirably well in such a task, particularly in moderate sample sizes. In these scenarios, $CvM_{n}$ performs slightly better than $KS_{n}$. Looking at the results from Shaikh2009's test, we note that bandwidth choices can play an important role, and the choice of the \textquotedblleft best\textquotedblright\ bandwidth $h_n$ via $c$ varies across DGPs. Perhaps, what is more important to emphasize in terms of power is that in all alternative hypotheses and sample sizes analyzed, our projection based tests have higher power than Shaikh2009's test, regardless of the bandwidth choice. Our proposed tests also dominate the balancing tests, as balancing tests have little to no power in all DGPs considered but $DGP2$, even with $ n=1,000$. Finally, the results in Table (ref) show that projection-based tests perform either better or as well as the other ICM tests in the DGPs considered. Among the considered DGPs, $DGP3$ is the only one where some existing specification testing has higher power than our proposed procedure. For this particular DGP, ICM type tests $KS^{trad}_n$ and $CvM^{trad}_n$ based on ((ref)) have higher power than our proposed tests when sample size $n$ is large. Given the discussion in Section (ref) and the fact that none of the ICM tests are strictly better than the others uniformly over the space of alternatives, such a finding does not come with a surprise.

Simulation 2

In this simulation, we push forward the dimensionality of the covariates to see how our proposed tests and the other alternative tests perform in scenarios with 10 continuous covariates. To investigate further this issue, we consider the following five DGPs:

align*[align* omitted — 524 chars of source]

where $X_{1}$, $X_{2}$ and $\varepsilon $ are defined as before, $\left\{ X_{i}\right\} _{k=3}^{10}$ are independent standard normal random variables, $D=1\left\{ D^{\ast }>0\right\} ,$ and $\varepsilon \protect\mathpalette{\protect\independentT}{\perp} X$, with $X=\left( 1,X_{1},X_{2},\dots ,X_{10}\right) ^{\prime }$. For each of these DGPs we consider the potential outcomes

equation*[equation* omitted — 143 chars of source]

where $m_{2}\left( X\right) =1+\sum_{j=1}^{10}X_{j}$, $u\left( 1\right) $ and $u\left( 0\right) $ are independent normal random variables with mean zero and variance $0.1$. The observed outcome is $Y=DY\left( 1\right) +\left( 1-D\right) Y\left( 0\right) $, and the true $ATE$ is 1.

For $DGP6-DGP10$, the $H_{0}$ considered is

equation[equation omitted — 240 chars of source]

We estimate $\theta _{0}$ by ML$.$ Note that $DGP6$ falls under $H_{0},$ whereas $DGP7$-$DGP10$ fall under $H_{1}$. The simulation results for $DGP6$- $DGP10$ are presented in Table (ref).

As before, we first discuss the size properties of the tests. From the results of $DGP6$, we find that $KS_{n}$ and $CvM_{n}$ tests are oversized when $n$ $=100$, but as sample size $n$ increases, the empirical size gets closer to its nominal value. The same holds true for ICM tests based on the unprojected process ((ref)). Shaikh2009's test tends to be conservative (with the exception when $c=0.01),$ and sensitive to the bandwidth choice. The traditional balancing tests, and the ICM tests based on ((ref)) are conservative, reflecting the \textquotedblleft curse of dimensionality\textquotedblright . Finally, note that when the propensity score is correctly specified, the finite sample properties of the $ATE$ estimator ((ref)) are good: the bias and the length of the 95% confidence interval get smaller when sample size increases, and the coverage probability is relatively close to its nominal value. When one trims observations with extreme estimated propensity score, these properties further improve; see Table (ref) in the Appendix B.

Note that when the propensity score is misspecified, the $ATE$ estimator ( (ref)) can be severely biased, and inference can be unreliable. Thus, tests with higher power to detect such misspecifications can prevent one to make misleading conclusions about the effectiveness of a given policy. What is clear from Table (ref) is that, regardless of the sample size and bandwidth considered, Shaikh2009's test seems to have little to no power to detect the alternatives described in $DGP7$, $DGP9$, and $DGP10$ . For $DGP8$, the maximum power for their test is approximately 55% when $ n=1,000$ and $c=0.1.$ However, with $c=0.01,$ the power of Shaikh2009 's test reduces to approximately 30%, highlighting again how important (and non-trivial) is to \textquotedblleft appropriately\textquotedblright\ choose the bandwidth. In sharp contrast with Shaikh2009's test, note that for moderately sized samples, our proposed $KS_{n}$ and $CvM_{n}$ tests have non-trivial power to detect all the alternatives. Our projection-based tests seem to dominate the other tests in the scenarios considered. Note that the traditional balancing tests have no power to detect the alternatives considered. Finally, ICM tests based on the traditional empirical process ( (ref)) have substantially less power than our proposed tests, reflecting the cost of the \textquotedblleft curse of dimensionality\textquotedblright . The power gains from using the projection-based process ((ref)) instead of the unprojected process ((ref)) can also be noted.

\afterpage{

landscape\begin{table}[htbp] \caption{Monte Carlo results under designs $DGP6$-$DGP10$} \begin{adjustbox}{ max width=1\linewidth, max totalheight=1\textheight, keepaspectratio} \begin{threeparttable} \begin{tabular}{ccccccccccccccccc} \hline \toprule DGP & $n$ & \multicolumn{1}{c}{$CvM_{n}$} & \multicolumn{1}{c}{$KS_{n}$} & \multicolumn{1}{c}{$T_{n}(0.01)$} & \multicolumn{1}{c}{$T_{n}(0.05)$} & \multicolumn{1}{c}{$T_{n}(0.10)$} & \multicolumn{1}{c}{$T_{n}(0.15)$} & \multicolumn{1}{l}{Max-$t$} & \multicolumn{1}{l}{$Wald$} & \multicolumn{1}{c}{$CvM_{n}^{unp}$} & \multicolumn{1}{c}{$KS_{n}^{unp}$} & \multicolumn{1}{c}{$CvM_{n}^{trad}$} & \multicolumn{1}{c}{$KS_{n}^{trad}$} & \multicolumn{1}{c}{Bias} & \multicolumn{1}{c}{CI length} & \multicolumn{1}{c}{Coverage} \\ \hline \midrule 6 & 100 & 8.50 & 9.40 & 5.80 & 2.50 & 1.00 & 0.20 & 1.10 & 1.60 & 7.80 & 9.40 & 1.10 & 1.70 & -0.29 & 3.71 & 89.40 \\ 6 & 200 & 5.70 & 6.10 & 4.00 & 2.50 & 1.00 & 0.40 & 0.90 & 0.30 & 5.70 & 6.50 & 2.30 & 3.90 & -0.08 & 2.10 & 91.00 \\ 6 & 400 & 5.60 & 6.00 & 5.10 & 2.40 & 0.90 & 0.30 & 2.80 & 0.70 & 5.80 & 6.30 & 3.30 & 4.50 & -0.02 & 1.41 & 91.90 \\ 6 & 600 & 5.00 & 5.20 & 4.90 & 1.70 & 0.60 & 0.20 & 3.40 & 1.60 & 3.70 & 5.00 & 3.80 & 3.70 & -0.02 & 1.11 & 90.00 \\ 6 & 800 & 5.40 & 5.20 & 5.10 & 2.70 & 1.40 & 0.70 & 5.10 & 3.40 & 4.60 & 5.00 & 2.60 & 4.20 & -0.01 & 0.97 & 93.30 \\ 6 & 1000 & 4.90 & 6.00 & 4.50 & 2.80 & 1.10 & 0.50 & 4.90 & 5.20 & 5.70 & 5.80 & 3.70 & 5.40 & -0.01 & 0.84 & 92.30 \\ \hline 7 & 100 & 9.50 & 10.60 & 4.10 & 1.60 & 0.70 & 0.00 & 0.60 & 5.30 & 6.80 & 8.40 & 2.50 & 3.80 & 0.22 & 6.97 & 96.60 \\ 7 & 200 & 15.40 & 14.90 & 5.10 & 3.30 & 2.00 & 1.10 & 0.70 & 1.20 & 13.60 & 13.90 & 2.70 & 6.00 & 0.59 & 4.43 & 99.10 \\ 7 & 400 & 26.40 & 24.20 & 5.40 & 6.60 & 6.80 & 5.20 & 0.60 & 0.30 & 23.00 & 19.90 & 4.60 & 7.90 & 0.61 & 2.66 & 99.10 \\ 7 & 600 & 37.20 & 33.00 & 9.10 & 10.70 & 11.10 & 8.70 & 0.10 & 0.00 & 33.70 & 30.10 & 8.30 & 9.30 & 0.56 & 1.97 & 96.60 \\ 7 & 800 & 49.50 & 46.50 & 9.90 & 16.70 & 17.00 & 13.90 & 0.10 & 0.00 & 46.40 & 37.70 & 13.80 & 13.70 & 0.59 & 1.69 & 89.30 \\ 7 & 1000 & 57.20 & 52.90 & 11.90 & 22.90 & 24.30 & 21.60 & 0.20 & 0.20 & 55.60 & 47.50 & 17.70 & 16.20 & 0.60 & 1.51 & 78.10 \\ \hline 8 & 100 & 9.50 & 10.60 & 5.70 & 1.30 & 0.80 & 0.20 & 2.00 & 14.30 & 8.20 & 9.90 & 2.20 & 3.70 & 0.61 & 10.68 & 96.20 \\ 8 & 200 & 18.30 & 18.70 & 4.40 & 4.20 & 3.60 & 2.80 & 1.00 & 4.40 & 14.80 & 15.50 & 3.40 & 5.30 & 1.12 & 6.58 & 99.40 \\ 8 & 400 & 46.50 & 41.80 & 8.10 & 16.00 & 15.80 & 11.60 & 0.50 & 0.90 & 41.70 & 37.10 & 9.10 & 9.10 & 1.36 & 4.44 & 98.50 \\ 8 & 600 & 66.30 & 62.10 & 15.80 & 27.50 & 28.10 & 24.20 & 0.20 & 0.50 & 60.60 & 54.20 & 20.80 & 16.10 & 1.33 & 3.42 & 88.30 \\ 8 & 800 & 79.90 & 74.40 & 23.30 & 42.70 & 45.60 & 40.00 & 0.50 & 0.60 & 74.90 & 68.30 & 30.90 & 22.30 & 1.34 & 2.91 & 66.00 \\ 8 & 1000 & 87.00 & 82.10 & 31.30 & 52.70 & 55.00 & 51.20 & 2.10 & 0.30 & 82.40 & 77.70 & 42.20 & 25.40 & 1.30 & 2.51 & 44.00 \\ \hline 9 & 100 & 10.00 & 12.00 & 3.50 & 2.40 & 0.40 & 0.20 & 1.70 & 6.80 & 10.10 & 12.40 & 1.20 & 2.50 & 0.06 & 7.76 & 88.30 \\ 9 & 200 & 15.10 & 15.10 & 4.70 & 3.70 & 1.80 & 0.80 & 1.60 & 4.10 & 12.40 & 12.00 & 4.00 & 7.10 & 0.60 & 5.02 & 94.60 \\ 9 & 400 & 30.40 & 26.80 & 4.50 & 3.10 & 2.80 & 2.90 & 0.80 & 2.70 & 21.30 & 17.00 & 8.50 & 12.00 & 0.97 & 3.70 & 97.30 \\ 9 & 600 & 44.10 & 36.10 & 5.70 & 6.10 & 6.70 & 6.00 & 0.50 & 3.30 & 32.50 & 25.30 & 17.10 & 20.10 & 1.00 & 3.01 & 89.20 \\ 9 & 800 & 57.60 & 50.20 & 6.60 & 8.80 & 11.00 & 10.40 & 0.70 & 3.00 & 46.80 & 35.60 & 25.40 & 29.30 & 1.03 & 2.63 & 78.60 \\ 9 & 1000 & 70.50 & 57.50 & 5.30 & 11.30 & 15.70 & 16.60 & 0.80 & 2.40 & 59.30 & 45.60 & 35.60 & 39.80 & 1.07 & 2.45 & 67.90 \\ \hline 10 & 100 & 7.50 & 8.30 & 4.10 & 1.30 & 0.60 & 0.30 & 0.00 & 0.00 & 6.80 & 8.70 & 4.00 & 3.90 & -0.09 & 2.37 & 97.30 \\ 10 & 200 & 10.00 & 10.10 & 5.10 & 1.80 & 0.80 & 0.30 & 0.00 & 0.00 & 10.30 & 9.30 & 4.40 & 4.90 & -0.10 & 1.26 & 96.50 \\ 10 & 400 & 19.20 & 16.70 & 5.20 & 3.80 & 2.80 & 1.80 & 0.00 & 0.00 & 20.60 & 14.90 & 6.90 & 8.60 & -0.10 & 0.81 & 93.60 \\ 10 & 600 & 36.20 & 28.10 & 4.60 & 7.30 & 6.80 & 4.80 & 0.10 & 0.00 & 37.70 & 27.70 & 10.10 & 8.20 & -0.11 & 0.65 & 91.40 \\ 10 & 800 & 50.20 & 38.90 & 7.70 & 11.30 & 13.20 & 11.10 & 0.10 & 0.00 & 52.80 & 38.80 & 15.20 & 12.60 & -0.12 & 0.55 & 87.50 \\ 10 & 1000 & 64.30 & 51.70 & 8.20 & 18.20 & 22.30 & 18.90 & 0.10 & 0.00 & 67.00 & 51.40 & 21.00 & 13.90 & -0.12 & 0.49 & 83.10 \\ \hline \bottomrule \end{tabular} \begin{tablenotes}[para,flushleft] { Note: Simulations based on 1,000 Monte Carlo experiments. “$CvM_{n}$” and “$KS_{n}$” stand for our proposed Cram\'{e}r-von Mises and Kolmogorov-Smirnov tests. “$T_{n}(c)$” stands for Shaikh2009's test, with bandwidth $h_{n}=cn^{-1/8}$. “Max-$t$” and “$Wald$” stand for, respectively, the Bonferroni-corrected Max-$t$-test, and Wald balancing tests based on $\widehat{ATE}_{n}\left( X^{j}\right)$ defined in ((ref)). “$CvM_{n}^{unp}$” and “$KS_{n}^{unp}$” are the Cram\'{e}r-von Mises and Kolmogorov-Smirnov tests based on the unprojected empirical process $\hat{R}_{n}(u)$ defined in ((ref)), whereas “$CvM_{n}^{trad}$” and “$KS_{n}^{trad}$” are defined analogously, but based on $\hat{R}_{n}^{trad}\left( x\right)$ as defined in ((ref)). Finally, “Bias”, “CI length', and “Coverage” stands for the average simulated bias, estimated 95% confidence interval length, and 95% coverage probability for the $ATE$ estimator $\widehat{ATE_{n}}$ as defined in ((ref)). All entries are proportions of rejections at 5% level, in percentage points, except “Bias” ,“CI length”, and “Coverage” (measure in percentage points), which are as described above. See the main text for further details.} \end{tablenotes} \end{threeparttable} \end{adjustbox} \end{table}

}

Overall, the simulation results highlight that the proposed projection-based tests perform favorably compared to other alternative testing procedures in terms of size and power. Importantly, the simulations illustrate that the gains in power can be credited to each distinguished feature of our tests, that is, $\left( i\right) $ the avoidance of smoothing parameters by using the ICM approach, $\left( ii\right) $ the dimension-reduction coming from considering $1\left\{ q\left( X,\theta _{0}\right) \leq u\right\} $ instead of $1\left\{ X\leq u\right\} $, and $\left( iii\right) $ the use of orthogonal projections. Given these attractive features, we believe that our tests can be of great use in practice.

Empirical Illustration

In this section, we provide an empirical illustration of our testing procedure. We revisit the analysis of Frankel2005 and Millimet2009, and study the effect of trade on the environment. More specifically, following Millimet2009, we assess the effect of a country being a member of the General Agreement on Tariffs and Trade (GATT) or World Trade Organization (WTO) on 5 different measures of environmental quality: per capita $CO_{2}$ emissions, average annual deforestation rate for 1990-1996, energy depletion, rural access to clear water and urban access to clear water.

As Millimet2009, we use three different covariates to model the probability of being a GATT/WTO member: real per capita GDP, land area per capita, and polity, which is a measure of how democratic (versus autocratic) is the structure of the government. The motivation to include these three covariates is to increase the plausibility of Assumption (ref) as discussed by Frankel2005. GDP per capita is associated with the probability of being member of the GATT/WTO and at the same time may have effects on different measures of environment quality, for instance, via the environmental Kuznetz curve. Land area per capita is another potential confounder since higher population density may lead to environmental degradation and \textquotedblleft larger\textquotedblright\ countries are more likely to trade more, affecting the probability of being member of he GATT/WTO. Finally, as noted by Frankel2005, low-democracy countries tend to have lower measures of environmental quality, and can also confound the effect of GATT/WTO membership. In what follows and in the same spirit of Millimet2009, we assume that Assumption (ref) holds after controlling for these three confounding factors\footnote{ If one finds the plausibility of this assumption rather low, all the estimates presented below should be interpreted as associations/correlations and not as causal effects. In light of Remark (ref), however, Assumption (ref) plays no role in our specification tests.}.

The unbalanced country-level panel data we use follows from Millimet2009, and includes observations from 1990 (before the WTO) and 1995 (after the WTO). However, it is important to have in mind that treatment is defined as being a GATT/WTO member, and therefore there are countries who were treated in both times, and others who were treated only in 1995. Table (ref) provides summary statistics and more detailed description of the variables. Finally, we highlight that the data we analyze is from an unbalanced country-level panel and instead of only considering the \textquotedblleft always observed\textquotedblright\ countries, we follow Millimet2009 and run a separate analysis for each outcome. For further details, see Section 4.1 of Millimet2009.

table[table omitted — 1,914 chars of source]

The main goal of this section is to assess the \textquotedblleft reliability\textquotedblright\ of different treatment effect measures by analyzing if different propensity score models are correctly specified or not. Given that the sample constitutes of an unbalanced panel, we estimate separate propensity score models for each outcome subsample. More specifically, for each outcome, we model the probability of a country being a GATT/WTO member ($D=1$ if member, $D=0$ otherwise) by a standard Probit model and consider two different specifications:

Spec1: $X$ includes real per capita GDP, land area per capita, and polity.

Spec2: $X$ is defined as in Spec1 but adds pairwise interaction terms between each covariate.

For each of these specifications, we test the null hypothesis

equation*[equation* omitted — 174 chars of source]

against $H_{1}$, which is simply the negation of $H_{0}$. Table (ref) reports the testing results for each specification, together with normalized IPW estimator for the $ATE$ based on ((ref)). We also consider the normalized IPW estimator for the $ATT$,

equation[equation omitted — 185 chars of source]

where $w_{1,i}^{treat}=D_{i}$, $w_{0,i}^{treat}=\left( 1-D_{i}\right) q\left( X_{i},\hat{\theta}_{n}\right) /\left( 1-q\left( X_{i},\hat{\theta} _{n}\right) \right) $, and $\bar{w}_{d,n}^{treat}$ is the sample mean of $ w_{d,i}^{treat},$ $d=\left\{ 0,1\right\} $. The associated standard errors and $p$-values are in parenthesis and brackets, respectively. Following\ Millimet2009, we trim observations with estimated propensity score outside the interval $[0.05,0.95]$ to avoid denominators arbitrarily close to zero. Bootstrapped $p$-values for our proposed specification tests are based on 100,000 bootstrap draws\footnote{ Note that the variables in Spec1 and\ Spec2 are all functions of the same three covariates: real per capita GDP, land area per capita, and polity. As so, Spec1 and Spec2 have the same information content with respect to the reliability of Assumption (ref).} .

At the 5% level we find that, based on the $CvM_{n}$ test statistic ((ref)), Spec1 is rejected for per capita $CO_{2},$ deforestation and energy depletion, but is not rejected for rural and urban accesses to clean water. The evidence of propensity score misspecification is weaker when using the $KS_{n}$ test statistic ((ref)). Spec2, on the other hand, is not rejected for any outcome at the usual significance levels, using either $CvM_{n}$ or $KS_{n}$ test statistic. Thus, our tests suggests that Spec2 should be preferred when analyzing per capita $ CO_{2},$ deforestation and energy depletion, whereas for urban and rural water access, our tests do not favor either specification.

table[table omitted — 2,790 chars of source]

Next we comment on the consequences of propensity score misspecification. For per capita $CO_{2}$, our results suggest that the overall effect of GATT/WTO membership on emissions is negative and statistically significant at the 5% level under both propensity score specifications. On the other hand, we find the effect of GATT/WTO membership on per capita $CO_{2}$ is not statistically significant among the treated sub-population using either specification. In terms of point estimates, however, there are important differences. For example, the $ATE$ point estimate under Spec1 (misspecified propensity score) is 30% higher (in absolute terms) than under Spec2. Note that the $0.3$ difference in $ATE$ represents roughly 8% of the overall per capita $CO_{2}$ emissions.

When we analyze the effect of GATT/WTO on deforestation and energy depletion, our results again highlight the consequences of propensity score misspecifications. We find that the $ATE$ point estimate for the effect of GATT/WTO membership on deforestation is 30% larger under Spec2 than under Spec1, and the $ATT$ point estimate for the effect of GATT/WTO membership on energy depletion is 25% smaller under Spec2 than under Spec1. Such large differences are economically significant, as the 0.08 difference in $ATEs$ on deforestation represents nearly 12% of the mean annual deforestation, and the 0.46 difference in $ ATTs$ on energy depletion represents nearly 15% of the mean energy depletion among countries in the sample. Interestingly, the $ATT$ on energy depletion is statistically significant at 10% level under Spec1, but not under Spec2, highlighting that propensity score misspecifications can also lead to invalid inference. Finally, we note that although our tests do not detect propensity score misspecification for the rural and urban water access, the results suggest that GATT/WTO membership is not statistically significant using either propensity score specification, perhaps because of the relatively high standard errors due to the limited sample size.

Overall, we find that propensity score misspecifications can affect both the economic and statistical conclusions about the effect of GATT/WTO membership on environmental quality. When the propensity score is correctly specified, our results suggest that GATT/WTO membership is associated with improved environmental performance in terms of $CO_{2}$ and energy depletion but not in terms of rural and urban water access. We also find that GATT/WTO membership is associated with higher deforestation, though the statistical evidence is relatively weaker. Furthermore, our results uncover interesting heterogeneity, as the aforementioned results are statistically significant for the overall population ($ATE$) but not for the treated-subpopulation ($ ATT$).

Conclusion

In this article, we have shown that, when propensity scores are correctly specified, a particular restriction between the propensity score CDFs of treated and control groups must hold. Based on such restriction, we propose new nonparametric projection-based tests for the correct specification of the propensity score. In contrast to other proposals, our tests are not severely affected by the \textquotedblleft curse of dimensionality\textquotedblright ', are not sensitive to the different estimation methods used to estimate the propensity score under the null, do not rely on the potentially ad hoc choice of bandwidths, and enjoy some optimal power properties against particular classes of alternative hypotheses. We have derived the asymptotic properties of the proposed tests, and have proved that they are able to detect local alternatives converging to the null at the parametric rate, and that critical values can be easily computed via a simple multiplier bootstrap procedure. Our Monte Carlo simulation study illustrates that, for a large class of alternatives, our projection-based tests perform better in finite samples than existing tests, though there are some classes of alternatives in which our tests have trivial power. All these finite sample findings are in line with our asymptotic results. Finally, our empirical application concerning the effect of trade on the environment shows the feasibility and appeal of our tests in relevant scenarios. Given that the validity of many policy evaluation procedures relies on the correct specification of the propensity score, we argue that the tests proposed in this article are important additions to the applied researchers' toolkit.

We would like to mention that, in general, our tests should be seen as a \textquotedblleft model validation\textquotedblright\ and not a \textquotedblleft model selection\textquotedblright\ procedure. Once a propensity score model is selected, our specification tests can provide evidence of its reliability or lack thereof. In case the putative propensity score model is rejected, one can consider more flexible specifications. For instance, one can add additional interaction terms into the original model, consider semiparametric single-index or partially linear models, among other possible strategies. Having said so, we emphasize that if one uses our testing procedure as a \textquotedblleft model selection\textquotedblright\ device, one must bear in mind that standard inference procedures for treatment effects may be invalid if one treats the resulting selected propensity score as the \textquotedblleft true\textquotedblright\ one, see e.g. Leeb2005. Thus, in case one uses our proposal for model selection, one must account for the model selection step in order to make valid inference about the treatment effect, see e.g. Belloni2014, Chernozhukov2016, and Belloni2017. A full discussion of this procedure is beyond the scope of this article and we leave it for future research.

Finally, we note that results in Lemma (ref) can also be used for estimating the propensity score such that, for a given specification $ q\left( X,\theta _{0}\right) $, condition $\mathbb{E}\left[ D-q\left( X,\theta _{0}\right) |q\left( X,\theta _{0}\right) \right] =0~a.s.$ is directly imposed. One could use the minimum distance method described in Dominguez2004 to estimate $\theta _{0}$, for example. A detailed discussion of this estimator is beyond the scope of this article and is deferred to future work.