Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,353 characters · 20 sections · 44 citation commands
Robust Estimation of Average Treatment Effects from Panel Data
The problem of estimating average treatment effects (ATE) from non-randomized studies has become quite intriguing to economists and social scientists in recent times. Although the term “treatment" was initially associated with the fields of medicine and agriculture, with the passage of time, it is being used in a more general sense where the effects of any specific experimentation, event or policy intervention are collectively termed as “treatment" effects. Impact of any such treatment is statistically estimated via the ATE across several domains of applications including, among others, economics, business and management sciences, social or political sciences, marketing, etc. For example, the ATE has been used to study the impact of sovereignty change and the CEPA implementation between mainland China and Hong Kong (starting in the first quarter of 2004) on the economic growth of Hong Kong in Hsiao/etc:2012. The impact of opening a new showroom of eyewear brand on its sales has been investigated in Li:2018. Further applications of the ATE can be found in, among many others, Bai/etc:2014,Ouyang/Peng:2015,Du/Zhang:2015.
As the name suggests, the ATE is intuitively computed by comparing the outcomes of units in the treatment group to those of units in the control group and is often expressed as the difference in their mean outcomes. But, such a simple approach can yield the true treatment effect only if the two groups are perfectly homogeneous in every possible aspect, which is often not the case in practice (not even in many planned randomized experiments if the sample sizes are not large enough). As put forward by Sakaguchi:2020, evaluation of policy (treatment) effects from non-randomized studies generally assumes that, given observed covariates, treatment assignments are independent of possible outcomes, which is often violated due to unobserved confounding variables. The main challenge to get correct treatment effect from such real-life (non-randomized) data, in fact, lies in estimating the counterfactual, the hypothetical outcome of the treated unit in the absence of treatment, with utmost accuracy via a proper adjustment for the confounding effects.
To outmanoeuvre the issue of unobserved confounding variables, suitable latent factor models are proposed to estimate the ATE from panel data containing multi-period observations. In any policy intervention (or event), which affects individuals or the society over time, panel data are the most common type available with some observations being before the intervention event and some being after it. So, it is important to develop appropriate methods for estimating ATE from panel data and, as expected, there are already various methods available for this purpose. Abadie:2005 proposed an innovative Difference-in-Differences (DID) approach for estimating the counterfactual with time-varying latent factors, which leads to a consistent estimator of the ATE even when the number of time periods is small provided the numbers of control and treatment units are large. However, there are some important limitations of this approach as it assumes no sample selection effect and that the average outcomes for the treatment and control units follow parallel paths over time in the absence of treatment. Li/Bell:2017a proposed the Augmented DID (ADID) approach for consistent estimation of the ATE using a linear regression of treatment units on the control units over the pre-treatment period; the resulting estimator is robust to the selection method for non-treated units and also performs well when the parallel path assumption is violated. For the cases exhibiting nonlinear and/or stochastic trends, Abadie/Gardeazabal:2003 proposed the synthetic control method (SCM) which uses a weighted average of control units to approximate the counterfactual outcome of the treated units in the absence of treatment; the weights are restricted to be non-negative and to sum to one. The performance of SCM relies crucially on the assumption that the weighted average of the control units' sample path is parallel to the treated unit's sample path in the absence of treatment. If this assumption is violated in practice, Doudchenko/Imbens:2016 suggested a modified SCM (MSCM) by relaxing the restriction that weights sum to one. This modification effectively relaxes the original parallel line assumption to a weaker condition that the weighted average of the control unit times a positive constant is parallel to the sample path of the treated unit in the absence of treatment. Note that the MSCM approach contains the SCM as a special case which, in turn, include the DID method for a particular choice of weights. The MSCM method performs well when the number of time periods (both pre and post) is large, but leads to larger estimation variance in the presence of a large number of control units. Hsiao/etc:2012 proposed a novel and flexible approach to estimate ATE which requires neither the assumption of no sample selection effect nor that of the parallel paths for treatment and control units in the absence of treatment. Furthermore, their method, which we will refer to as the HCW method, yields consistent estimate of the ATE that is additionally robust to any nonlinear functional form. However, similar to the SCM, having a large number of control units can lead the HCW method to over-fit in-sample data and generate imprecise out-of-sample predictions.
A major concern for all the above methods of estimating ATE from panel data is their extreme non-robustness against data contamination (e.g., outliers), as they are all averaging and/or least-squares techniques. Since outliers are not infrequent in real-life datasets, these existing procedures lead to non-robust and incorrect inference about the ATE in the presence of data contamination. When a treatment intervention is given to a particular unit, one may carefully collect its response during the follow-up (post-treatment) period to avoid any noise; but it is not always feasible to protect the observations from contamination on all the control units (and also for the treated unit in pre-treatment time-points). So, a robust procedure that can yield stable inference about the ATE from (possibly noisy) panel data is extremely necessary for correct policy planning and other insight generation in practical applications. The present paper solves this crucial issue by proposing a new robust estimation and subsequent inference procedures for the ATE under panel data set-up, leading to stable and highly efficient insights even under data contamination.
Among several existing approaches of robust parameter estimation, we consider the minimum distance methodology with the popular density power divergence (DPD) of Basu/etc:1998. This minimum DPD estimation has become extremely popular in recent times due to its easy interpretation as a generalization of the maximum likelihood estimator (MLE) and high asymptotic efficiency along with the desired robustness properties Basu/etc:2011. As a result, the minimum DPD estimator (MDPDE) has also been extended to several advanced data types and inference problems; see, e.g., Ghosh/Basu:2013,Ghosh/Basu:2015,Ghosh/Basu:2016,Ghosh/Basu:2016b,Durio/Isaia:2011. For the sake of completeness, a brief review of the MDPDE is provided in Appendix (ref). In this paper, we particularly use the MDPDE, as defined for the linear regression in Durio/Isaia:2011, Ghosh/Basu:2013, along with the HCW factor model for panel data set-up to develop a new robust and efficient ATE estimator. The nice asymptotic properties of the MDPDE would translate to our proposed estimate of ATE establishing its consistency and asymptotic normality. Although the use of MDPDE in constructing the robust ATE estimator is somewhat straightforward, derivations of the asymptotic properties of the resulting ATE estimator are rather nontrivial which are the major contributions of the present paper along with their innovative applications in (robustly) studying treatment effects. Using the proposed ATE estimates, we also develop robust testing procedures for both one and two sample hypotheses involving ATE. The claimed robustness of the proposed estimation and testing procedures, against data contamination in the pre-treatment periods, would be illustrated through appropriate Monte-Carlo simulation studies and also theoretically via the influence function analysis of the ATE estimator. An appropriate modification is also suggested to make the proposed ATE estimator robust against contamination in the post-treatment period's data, with empirical illustrations. Finally, our proposal is applied to examine the long-term effects of the 2004 Indian Ocean earthquake and tsunami on the economic growth (measured via GDP) of five countries, namely Indonesia, Sri Lanka, Thailand, India and Maldives, that had received the major instant shock.
The rest of the paper is structured as follows. In Section (ref) we discuss the model setup and the proposed robust ATE estimator, while its asymptotic properties are derived in Section (ref). Statistical tests based on the new ATE estimator are presented in Section (ref). Section (ref) presents simulation studies to compare the efficacy of our proposed ATE estimators over the estimating ones and also to illustrate the finite sample performances of the associated robust tests. The (local) robustness of our ATE estimator is further studied theoretically in Section (ref) through the influence function analysis, where another modified version is discussed to gain robustness against data contamination in the post-treatment period. Section (ref) discusses the real data applications and some concluding remarks are presented in Section (ref). For the brevity in presentation, some proofs are moved to Appendices (ref) and (ref).
Suppose that we have a panel dataset consisting of observations on $N$ units over $T$ time periods, and let $y_{it}$ denote the response of $i$-th unit at $t$-th time point for $i=1, \ldots, N$ and $t=1, \ldots, T$. To identify the counterfactuals, let us also denote by $y_{it}^1$ and $y_{it}^0$ the outcome of unit $i$ in period $t$ with and without treatment, respectively. Then the treatment effect of the $i^{th}$ unit at a post-treatment time-point $t$ is defined as
However, for any $i$, $y_{it}^0$ and $y_{it}^1$ are not simultaneously observed at a time-point ($t$). Thus, the observed data actually have the form $y_{it}=d_{it}y_{it}^1+(1-d_{it})y_{it}^0$ for each $i$ and $t$, where
The difficulty in estimating the counterfactual outcome $y_{it}^0$, when $d_{it}=1$, makes the estimation of ATE cumbersome. For estimation of the counterfactual outcome of the treated unit in the post-treatment period, we consider the HCW factor model Hsiao/etc:2012 having several nice properties as described in the introduction (Section (ref)). It assumes the existence of $K$-dimensional unobservable (latent) factors $\boldsymbol{f}_t$ for every time point $t$ such that
where $a_i$ is a unit specific intercept, $\boldsymbol{b}_i$ is the $K$-dimensional factor loadings for $i$-th unit associated with the factor $\boldsymbol{f}_t$ and $u_{it}$ is a zero mean, weakly dependent and weakly stationary error term. The above relation can also be written in a matrix form as
where $\boldsymbol{y}_t^0= (y_{1t}^0,...y_{Nt}^0)'$, $ \boldsymbol{a} =(a_1,.....a_N)'$, $\boldsymbol{B}= [\boldsymbol{b}_1, ..., \boldsymbol{b}_N]'$ and $\boldsymbol{u}_t = (u_{1t},...u_{Nt})'$.
Now, let us assume that one unit is exposed with the treatment at time $T_1$ $(0< T_1 < T-1)$; without loss of generality, suppose that the first unit ($i=1$) is the treated one to avoid complex notation. Then, the responses of the first unit must satisfy the relations
where the treatment effect $\Delta_{1t}$ is as defined in ((ref)) for $i=1$. Thus, once we obtain an estimate $\widehat{y_{1t}^0}$ of the counterfactual $y_{1t}^0$ for a post-treatment time point $t$, the treatment effect at that time-point can be easily estimated as $\widehat{\Delta}_{1t} = y_{1t}-\widehat{y_{1t}^0}$; these are finally averaged over all the post treatment periods to get an estimate of the ATE.
However, the latent factors in ((ref)) are not observable and hence we cannot directly use it to estimate the counterfactual $y_{1t}^0$. An indirect way was proposed in Hsiao/etc:2012 which suggests to transform Equation ((ref)) as follows. For any $r$-vector $\boldsymbol{v}$, let us denote the $(r-1)$-vector obtained by removing the first element of $\boldsymbol{v}$ as $\widetilde{\boldsymbol{v}}$. Now, choosing an appropriate vector $\boldsymbol{\gamma}\in\mathcal{N}(\boldsymbol{B}')$, the null space of $\boldsymbol{B}'$, with the first element of $\boldsymbol{\gamma}$ being unity (with suitable scaling), and pre-multiplying ((ref)) by ${\boldsymbol{\gamma}}'$, we get
where $\delta_1={\boldsymbol{\gamma}}'\boldsymbol{a}$, $\boldsymbol{\delta}=\widetilde{\boldsymbol{\gamma}}'$, $\eta_{1t}=\boldsymbol{\gamma}'\boldsymbol{u_t}$ and $\widetilde{\boldsymbol{y}_t^0}=(y_{2t}^0,\cdots,y_{Nt}^0)$. Note that, Equation ((ref)) holds for every $t=1, \ldots, T$ and justify the suggestion of Hsiao/etc:2012 that $\widetilde{\boldsymbol{y}_t^0}$ can be used in lieu of $\boldsymbol{f}_t $ to predict $y_{1t}^0$. In particular, at the pre-treatment time periods we have $y_{it}^0=y_{it}$ so that ((ref)) reduces to
The regression coefficients $\delta_1$ and $\boldsymbol{\delta}$ are then estimated by the least square method from ((ref)). If these estimates are denoted by $\widehat{\delta_1}$ and $\widehat{\boldsymbol{\delta}}$, then the counterfactual outcome $y_{1t}^0$, for the post-treatment period, can be estimated via $\widehat{y_{1t}^0}=\widehat{\delta_1}+\widehat{\boldsymbol{\delta}}'\widetilde{\boldsymbol{y}_t}$ for $t= T_1+1,...T$, since $y_{it}^0=y_{it}$ for all $t$ when $i\neq1$.
At this point, we should note that some assumptions are required for the derivation of ((ref)) to go through as noted in Hsiao/etc:2012; however, all the assumptions of Hsiao/etc:2012 were indeed not necessary as shown later in Li/Bell:2017b. Accordingly, let us note down the following set of sufficient conditions.
Note that only (A2) is needed for the calculations leading to ((ref)). Additionally, (A1) was used in Li/Bell:2017b to show that the error term $\eta_{1t}$ and the coefficients $\delta_1$, $\boldsymbol{\delta}$ can be chosen (with some translation if required) so as to satisfy $E(\eta_{1t})=0$ and $E(\eta_{it}\widetilde{\boldsymbol{y}_t})=0$ for all $t$, which ensure the consistency of the least-square based estimate of the ATE $\Delta_1$ along with (A3). Note also that, intuitively $\Delta_{1t}$ should be assumed to be independent of $\{\widetilde{\boldsymbol{y}_t}\}$ for any $t=1, \ldots, T$, since the treatment effect should not depend on the control units.
We now discuss our proposed approach for the robust estimation of ATE from panel data. Let us consider the HCW factor model set-up and notation from the previous subsection so that the transformed Equations ((ref)) and ((ref)) are valid (requires Assumption (A2)). As mentioned before, we will use the popular minimum DPD estimation approach to get robust estimates of the parameters $\delta_1$, $\boldsymbol{\delta}$ and use them to derive the robust ATE estimator. In order to achieve greater efficiency, this approach of parametric robustness considers outliers defined with respect to a given model distribution that the majority of observations follow. Thus, in order to get such a parametric robust estimate of the regression coefficients from ((ref)), we first consider a model density for the errors $\eta_{1t}$. So, instead of Assumption (A1), we will assume that these error terms, after suitable adjustment through linear projection as in Li/Bell:2017b, satisfy the following condition for a tuning parameter $\alpha\geq 0$.
Under Assumption (A4), it follows from Equation ((ref)) that $y_{1t}$ has density $f_t(y_{1t}) = \frac{1}{\sigma}f\left(\frac{y_{1t}-\delta_1 -\boldsymbol{\delta}^{\prime}\widetilde{\boldsymbol{y}_t}}{\sigma}\right)$ for each $t=1, \ldots, T_1$ and they are independent given $\widetilde{\boldsymbol{y}_t}$. Thus, as detailed in Appendix (ref), the MDPDE of the model parameters $\boldsymbol{\theta}=\left(\delta_1, \boldsymbol{\delta}, \sigma^2 \right)$, at any given $\alpha>0$, can be obtained by minimizing the objective function ((ref)) over $t=1, \ldots, T_1$ with $f_i \equiv f_t$, which simplifies to the form
In the particular case of normal error distribution, i.e., when $f$ is the standard normal density, the above MDPDE objective function can be further simplified as
which we need to minimize with respect to $\boldsymbol{\theta}$ to obtain the MDPDEs of the regression coefficients and the error variance. Note that, unlike the least squared based methods, the MDPDE objective function $H_{T_1}^{(\alpha)}(\boldsymbol{\theta})$ in ((ref)), or its simpler version in ((ref)), does not have an explicit minimizer. So, we need to compute the MDPDE by numerically minimizing $H_{T_1}^{(\alpha)}(\boldsymbol{\theta})$ via an appropriate optimization technique.
Equivalently, the MDPDEs can also be obtained by numerically solving the corresponding estimating equations obtained by equating the derivative of $H_{T_1}^{(\alpha)}(\boldsymbol{\theta})$, with respect to $\boldsymbol{\theta}$, to zero. For the special case of normal error density, by differentiating $H_{T_1}^{(\alpha)}(\boldsymbol{\theta})$ in ((ref)) with respect to $\delta_1$, $\delta_j$, $j=2,3,\cdots N$ and $\sigma$, we get the corresponding MDPDE estimating equations as given by
For the general error distribution $f$, assuming $f$ to be differentiable and letting $u=f'/f$, the general set of MDPDE estimating equations can be derived as
Let us now denote the MDPDE of $\boldsymbol{\theta}=\left(\delta_1, \boldsymbol{\delta}, \sigma^2 \right)$, obtained with tuning parameter $\alpha> 0$, as $\widehat{\boldsymbol{\theta}}_\alpha=\left(\widehat{\delta}_{1,\alpha}, \widehat{\boldsymbol{\delta}}_\alpha, \widehat{\sigma}_\alpha^2 \right)$. Using them, we can (robustly) estimate the counterfactual outcome $y_{1t}^0$ at any post-treatment time-period as
and the corresponding treatment effects are estimated by
Finally, a robust estimate of the ATE $\Delta_1$ is given by
Due to the use of the MDPDE of the regression coefficients $(\delta_1, \boldsymbol{\delta})$ and the averaging at the final step, we will refer to the estimator $\widehat{\Delta}_1^{(\alpha)} $ as the Mean-MDPDE of the ATE $\Delta_1$ with tuning parameter $\alpha>0$.
Let us first discuss the large sample behavior of the MDPDE of the regression coefficients and the error variance following Ghosh/Basu:2013. In this context, we need some additional assumptions as follows.
It is easy to see that Assumption (A5) holds if $\{\boldsymbol{x}_t\}$ is a weakly dependent and weakly stationary process with $\boldsymbol{\Sigma}_x = \displaystyle {\mbox{plim}}_{T_1\rightarrow\infty }T_1^{-1}(\boldsymbol{X}_0'\boldsymbol{X}_0)$ being invertible. Then, the asymptotic properties of the MDPDEs $\widehat{\boldsymbol{\beta}}_\alpha=\left(\widehat{\delta}_{1,\alpha}, \widehat{\boldsymbol{\delta}}_\alpha' \right)$ and $\widehat{\sigma}_\alpha^2$ of the regression parameters $\boldsymbol{\beta}=\left(\delta_1, \widetilde{\delta}' \right)$ and $\sigma^2$, respectively, can be obtained by an application of the general theory from Ghosh/Basu:2013. Specifically, under Assumptions (A4)--(A5), we can show the following result for any $\alpha\geq 0$; see Appendix (ref) for proof.
In particular, when the error density $f$ is standard normal, the conditions of Corollary (ref) hold true and the asymptotic variances of the MDPDEs $\widehat{\boldsymbol{\beta}}_\alpha$ and $\widehat{\sigma}_\alpha^2$ can be further simplified with
These asymptotic distributional results of the MDPDEs will be used, in the following subsection, to derive the consistency and asymptotic distribution of the proposed Mean-MDPDE estimator $\widehat{\Delta}_1^{(\alpha)}$ of the ATE $\Delta_{1}$ as $T_1, T_2 \rightarrow\infty$. For simplicity, throughout the rest of the paper, we will assume that the conditions of Corollary (ref) hold so that the asymptotic distributions of $\widehat{\boldsymbol{\beta}}_\alpha$ and $\widehat{\sigma}_\alpha^2$ are independent; all the results, however, can easily be extended for the general case using Theorem (ref) with messier formulas and notations.
In order to derive the asymptotic properties of the ATE estimator $\widehat{\Delta}_1^{(\alpha)} $, let us first note its decomposition as given by
After some algebra, we get
Note that, by Theorem (ref), $\left(\widehat{\boldsymbol{\beta}}_\alpha - \boldsymbol{\beta}\right) = O_p(T_1^{-1/2})$ as $T_1\rightarrow\infty$. Further, via the laws of large numbers, we get
Therefore, the first term in ((ref)) is $O_p(T_1^{-1/2})$ as $T_1, T_2\rightarrow\infty$. Additionally, due to Assumption (A3), we have $T_2^{-1/2}\sum_{t=T_1+1}^{T}\left(\Delta_{1t} - \Delta_1 + \eta_{1t} \right)=O_p(1)$. Hence, combining all these, we get the consistency of the Mean-MDPDE $\widehat{\Delta}_1^{(\alpha)}$ as presented in the following theorem.
Next, in order to derive the asymptotic distribution of the ATE estimator $\widehat{\Delta}_1^{(\alpha)}$, we need appropriate assumption to make the two terms in ((ref)) asymptotically independent. In this respect, we consider the following general assumption.
Note that Assumption (A6) can be shown to hold in several practical scenarios under suitable mixing conditions for the processes $\{\boldsymbol{x}_t, \eta_{1t}\}$ for the whole time-period $t=1, \ldots, T$. Such a mixing condition is also used in Li/Bell:2017b while deriving the asymptotic normality of their ATE estimator. Particularly, (A6) holds if the variables from the pre-treatment period are independent of their values in the post-treatment period, e.g., if $(\boldsymbol{x}_t, \eta_{1t})$s are independently distributed over $t$. With this additional assumption, we can now derive the asymptotic distribution of $\widehat{\Delta}_1^{(\alpha)}$ by computing the asymptotic variances of the two terms in ((ref)) individually, which is presented in the following theorem.
Proof: Let us start from Equation ((ref)) to deduce $\sqrt{T_2}\left(\widehat{\Delta}_1^{(\alpha)} - \Delta_{1} \right) = S_1 + S_2$, where
By Assumption (A3), we directly get that $S$ is asymptotically normal with mean zero and variance ${\Sigma}_2$ as $T_2 \rightarrow\infty$. . Therefore, the theorem will follow in view of Assumption (A6) if we can show that $S_1$ is asymptotically normal with mean zero and variance $\boldsymbol{\Sigma}_1 = \omega v_\beta(\alpha) \sigma^2 E[\boldsymbol{x}_t]'\boldsymbol{\Sigma}_xE[\boldsymbol{x}_t] $ as $T_1, T_2 \rightarrow\infty$.
But, by Theorem (ref), we have that $\sqrt{T_1}\left(\widehat{\boldsymbol{\beta}}_\alpha - \boldsymbol{\beta}\right)$ is asymptotically normal with mean zero and variance $v_\beta(\alpha)\sigma^2\boldsymbol{\Sigma}_x^{-1}$ and that $\left[\frac{1}{T_2}\sum_{t=T_1+1}^{T}\boldsymbol{x}_t\right]$ converges in probability to $E[\boldsymbol{x}_t]$, viz. Eq. ((ref)), as $T_1, T_2 \rightarrow\infty$. Combining them via Slutsky's theorem and Assumption (A6), we get the asymptotic normality of $S_1$ with the asymptotic mean zero and variance $\left(\sqrt{\omega} E[\boldsymbol{x}_t]\right)'[v_\beta(\alpha)\boldsymbol{\Sigma}_x]\left(\sqrt{\omega} E[\boldsymbol{x}_t]\right) = \boldsymbol{\Sigma}_1$. This completes the proof. {$\square$}
It is important to note that the asymptotic variance of the Mean-MDPDE of ATE depends on the tuning parameter $\alpha$ only through the term $v_\beta(\alpha)$. For the normal error distribution, $v_\beta(0)=1$ and hence the asymptotic variance $\Sigma(\alpha=0) $ coincides with the asymptotic variance of the HCW estimator as given in Theorem 3.2 of Li/Bell:2017b. As $\alpha>0$ increases, the values of $v_\beta(\alpha)$ increases leading to an increase in the asymptotic variance of the proposed ATE estimator, a price in order to achieve robustness under data contamination. However, by examining the values of $v_\beta(\alpha)$, e.g., in Figure (ref), it can be seen that the loss in efficiency (in terms of the inverse of the asymptotic variance) is not quite significant at small values of the tuning parameter $\alpha>0$. This fact that $\alpha$ provides a trade-off between robustness and efficiency of the proposed ATE estimator is consistent with the literature of MDPDE Basu/etc:2011 and, in general, the notion of parametric robustness Hampel/etc:1986.
Once we have an estimate of the ATE, we need to compute its standard error in order to have an idea about the credibility of our ATE estimate. For this purpose, we can use Theorem (ref) to get an asymptotic standard error for the proposed Mean-MDPDE of the ATE by consistently estimating the asymptotic variance $\Sigma(\alpha)$. If $\widehat{\Sigma}(\alpha)$ denote its consistent estimator, then the estimated asymptotic standard error of the ATE estimator $\widehat{\Delta}_1^{(\alpha)}$ is given by $\sqrt{\widehat{\Sigma}(\alpha)/T_2}$.
Now, in order to get the consistent estimate of $\Sigma(\alpha)$, we note that its second component ${\Sigma}_2$ is defined in Assumption (A3) in exactly the same way as in Li/Bell:2017b. Hence, following the arguments given in Li/Bell:2017b, a consistent estimator of ${\Sigma}_2$ is given by
where $l$ is chosen in such a way that $l\rightarrow\infty$ and $l/T_2 \rightarrow 0$ as $T_2 \rightarrow\infty$ (e.g., $l=O(T_2^{1/4})$, see also Newey/West:1987,White:2001). However, if $\Delta_{1t}$ and $\eta_{1t}$ are serially uncorrelated, then an alternative simpler consistent estimate of ${\Sigma}_2$ can be obtained as Li/Bell:2017b
Next, in order to consistently estimate the first component of $\Sigma$, we note that the term $\boldsymbol{\Sigma}_x$ can be consistently estimated by $T_1^{-1}(\boldsymbol{X}_0'\boldsymbol{X}_0)$. Similarly, a consistent estimate of $E[\boldsymbol{x}_t]$ is given by $\left[\frac{1}{T_2}\sum_{t=T_1+1}^{T}\boldsymbol{x}_t\right]$. On the other hand, the MDPDE $\widehat{\sigma}_\alpha^2$ is consistent for $\sigma^2$ in view of Theorem (ref). So, estimating $\omega$ by $T_2/T_1$, a consistent estimate of the first term of $\Sigma(\alpha)$ is then given by $$ v_\beta(\alpha) ~ \frac{T_2}{T_1} ~\widehat{\sigma}_\alpha^2 \left[\frac{1}{T_2}\sum_{t=T_1+1}^{T}\boldsymbol{x}_t\right]' T_1(\boldsymbol{X}_0'\boldsymbol{X}_0)^{-1}\left[\frac{1}{T_2}\sum_{t=T_1+1}^{T}\boldsymbol{x}_t\right]. $$ Upon simplification, the final consistent estimate of the asymptotic variance $\Sigma(\alpha)$ is given by
where $\widehat{{\Sigma}_2}$ can be taken as in ((ref)) or ((ref)) depending on our assumption.
In this section, we define a class of robust tests for hypothesis regarding the ATE, using the proposed Mean-MDPDE of the ATE, so that these tests remain stable under arbitrary small departures from the specified (parametric) model distributions. Proofs of all the results of this section are given in Appendix (ref).
Consider first the simplest problem of testing the simple null hypothesis $H_0: \Delta_1=\Delta_{10}$, where $\Delta_{10} \in \Theta \subseteq \mathbb{R}$. Although tests based on MLEs have several nice optimum properties, they are highly non-robust in presence of outliers in the sample. To circumvent this issue of non-robustness, we base our test statistic on the proposed Mean-MDPDE $\widehat{\Delta}_1^{(\alpha)}$ of the ATE, with tuning parameter $\alpha>0$. Noting that $\sqrt{\widehat{\Sigma}(\alpha)/T_2} $ is the standard error of the Mean-MDPDE, we define the one-sample test statistic for testing $H_0:\Delta_1=\Delta_{10}$ as given by
Using the asymptotic properties of the Mean-MDPDE presented in Section (ref), we can easily see that the asymptotic null distribution of the proposed test statistic $W_1^{(\alpha)}$ is indeed $\mathcal{N}(0,1)$. So, denoting $z_\tau$ to be the upper $100(1-\tau)$-th quantile of the standard normal distribution, our rejection regions at the level of significance $\tau$, against different alternative hypotheses, would be as follow.
Next we shall derive an approximate power function $\beta_{W_1}^{(\alpha)}$ of the proposed test statistic in the following theorem which would, in turn, prove the consistency of the proposed test. In the following, we present results only for the first alternative $H_1: \Delta_1>\Delta_{10}$; similar results can also be derived for the other two cases.
In order to produce a nontrivial asymptotic power, we can consider contiguous alternative hypotheses given by
where $d \in \mathbb{R}\setminus\{0\}$ such that $\Delta_{1,T_2} \in \Theta$. The following theorem presents the asymptotic distribution of the test statistics under $H_{1,T_2}$ and subsequently the expression of asymptotic contiguous power.
The problem of comparing ATEs arises when we want to study the impact of a social policy on two independent populations. Examples may include the impact of government education policy among school-going children in urban and rural sectors. The primary interest in such cases is to test whether the difference between the two ATEs is statistically significant or not, i.e., we want to test the two-sample hypothesis $H_0:\Delta_{11}=\Delta_{12}$, where $\Delta_{1j}\in \mathbb{R}$ denote the ATE in the $j$-th population for $j=1,2$. We use their corresponding Mean-MDPDEs to construct a robust test for $H_0$.
Let $T_{1j}$ and $T_{2j}$ be the number of observations available in pre- and post-treatment periods, respectively, for the $j$-th population and $\widehat{\Delta}_{1j}^{(\alpha)}$ be the Mean-MDPDE of the ATE $\Delta_{1j}$ with tuning parameter $\alpha>0$ for each $j=1, 2$. Assuming identifiability of the model family, the difference between the two estimators $\widehat{\Delta}_{11}^{(\alpha)}$ and $\widehat{\Delta}_{12}^{(\alpha)}$ gives us an idea of the difference between the ATEs in the two samples and, hence, indicate any departure from the null hypothesis. Denoting the standard error of $\widehat{\Delta}_{1j}^{(\alpha)}$ by $\sqrt{\widehat{\boldsymbol{\Sigma}_j}(\alpha)/T_{2j}}$ for $j=1, 2$, the two-sample robust test statistic for testing $H_0:\Delta_{11}=\Delta_{12}$ can be defined as
If the assumptions of Theorem (ref) hold for both the populations, then the two Mean-MDPDE estimators are asymptotically normally distributed and, as a consequence, one can easily show that the asymptotic distribution of $W_2^{(\alpha)}$ under $H_0:\Delta_{11}=\Delta_{12}$ is indeed standard normal. So, we reject $H_0$ against $H_1: \Delta_{11}>\Delta_{12}$, at a level of significance $\tau$, if $W_2^{(\alpha)}>z_\tau$. Tests for other alternatives can also be constructed in a routine manner.
As in the case of one sample problem, we can also derive an approximate power function $\beta_{W_2}^{(\alpha)}$ of the proposed two-sample test based on $W_2^{(\alpha)}$, which is presented in the following theorem.
We can again see from Theorem (ref) that $$ \lim_{T_{21},T_{22}\to +\infty}\beta_{W_2}^{(\alpha)}(\Delta^*_{11},\Delta^*_{12})= 1, ~~~~~\mbox{ for any } ~\Delta^*_{11} > \Delta^*_{12}, ~~\alpha>0. $$ Therefore, the proposed two-sample test based on $W_2^{(\alpha)}$ is also asymptotically consistent. We can, however, study its nontrivial asymptotic power by considering contiguous alternative hypotheses of a general form given by
We have conducted extensive simulation studies to examine the finite-sample performances of our proposed Mean-MDPDE of the ATE by examining their empirical bias and mean-square error (MSE) under both uncontaminated (pure) as well as contaminated data. Additionally, for comparison purpose, we also compute the same error measures for existing common ATE estimators, namely the HCW estimator Hsiao/etc:2012, the DID estimator Li/Bell:2017a, the augmented DID (ADID) estimator Li/Bell:2017a and the MSCM estimator Li:2020 which are then compared with our proposal in terms of robustness and efficiency. \\
\noindentSimulation Set-up: We consider the simulation experiments following Li/Bell:2017b, where one treatment and two control units ($N=3$) are observed over $T$ time-points and the first unit got the treatment at time-point $T_1+1$. To generate the observations according to the HCW factor model, we first simulate one unobserved latent variable ($K=1$) from an AR(1) process as $f_t$ = $0.5f_{t-1} + v_t$ with $v_t$ being independently distributed as $N(0,1)$ for each $t=1, \ldots, T$. Then the observations of the three units are generated via $\boldsymbol{y}_t= (y_{1t}, y_{2t},y_{3t})' = \textbf{a} + \textbf{b} f_t + \textbf{u}_t$, for every $t= 1,...,T$, where $\textbf{a}= (1,1,1)'$, $\textbf{b}=(1,1,1)'$ and $\textbf{u}_t= (u_{1t}, u_{2t},u_{3t})'$ with $u_{jt}\sim \mathcal{N}(0,1)$ independently for each $j=1,2,3$. To incorporate the treatment effects on the first unit, we modify it as $y_{1t} \leftarrow y_{1t} + \Delta_{1t}$ for the post-treatment time points $t=T_1+1, \ldots, T_2$, where $\Delta_{1t}$s are the individual treatment effects at time $t$ and are generated by $\Delta_{1t}= \exp(z_t)/[1+\exp(z_t)]+1$ with $z_t=0.5z_{t-1}+\epsilon_t$ and $\epsilon_t\sim \mathcal{N}(0,0.25)$. These errors ($\epsilon_t$s) are generated independent of each other and also independent of $(\widetilde{y}_t, u_{jt})$ for all $t= 1,...T$ and $j=1,2,3$. Then, we estimate the ATE from the sample data (using different estimators) and compute the empirical bias and MSE based on 2500 replications and against the true ATE value $\bar{\Delta}_{1}=T_2^{-1}\sum_{t=T_1+1}^{T}\Delta_{1t}$.
Additionally, in order to study the robustness of the estimators, we repeat the above simulation exercises by contaminating 5% or 20% of the sample observations on the first unit (treated response). These contaminated observations (randomly selected from the pool) are generated by using $u_{1t} \sim \mathcal{N}(5,1)$ instead of $\mathcal{N}(0,1)$ either in the pre-treatment or in the post-treatment periods. \\
\noindentSimulation Results: The resulting bias and MSE values of different ATE estimators are presented in Tables (ref) and (ref), respectively, for $T_1= 100,200,400$ and $T_2 = T-T_1 = 20,80,320$. It is clearly seen that that, when there is no data contamination, the DID estimator has the least MSE followed by HCW estimator. But, our proposed Mean-MDPDE at smaller $\alpha>0$ are also quite competitive; their loss in efficiency (in terms of MSE) is not quite significant and further diminishes as the sample size increases. On the contrary, in the presence of contamination in pre-treatment periods, the proposed Mean-MDPD estimator(s) clearly outsmarts all the existing ATE estimators; in this respect, Mean-MDPDE with larger $\alpha>0$ performs better indicating their increasing robustness properties with increase in $\alpha$.
However, the mean-MDPDE fails to generate robust ATE estimators in the presence of contamination in post-treatment periods, when it performs mostly at par with all existing estimators. Such robustness behavior of our Mean-MDPDE of the ATE will be further justified theoretically via the influence function analyses in Section (ref).
We also repeat the same simulation study as described in the previous subsection to investigate the performances of the proposed robust Wald-type tests for testing significance of the ATE for different finite sample-sizes $(T_1, T_2)$. In particular, we test for the hypothesis $H_0:$ $\Delta_1=0$ against $H_1:\Delta_1 \neq 0$ and plot the empirical power functions, obtained based on $5000$ replications at the 5% level of significance, for pure data as well as 10% contaminated data in Figure (ref). The contamination is taken in the pre-treatment period as in the previous subsection. Note that the value of the power function at $\Delta_1=0$ yields the empirical size of the corresponding tests. For the sake of comparisons, the power function for the Wald-type test based on the HCW estimator is also plotted in the same Figure (ref).
It can be seen from the figures that the robust MDPDE-based Wald type tests perform quite similarly to the classical test under pure data; they all maintain the empirical size at the specified level and have powers increasing to one as the sample size increases or as the alternative moves away from the null (i.e., $\Delta_1$ moves away from zero). Under contamination, however, the classical HCW estimator becomes extremely non-robust and hence the corresponding test have significantly inflated size and decreased power at any fixed alternative ($\Delta_1>0$) compared to the pure data case. On the contrary, the proposed MDPDE-based tests with higher values of $\alpha>0$ remain stable, both in terms of size and power, even under contamination indicating their claimed robustness benefits.
{\color{black}
}
It is important to note that the proposed ATE estimator and its performance crucially depend on a tuning parameter $\alpha$ $(0\leq \alpha \leq 1)$. By varying this tuning parameter, the strength of the estimator can be controlled via a trade-offs between efficiency and robustness of the resulting inference. Choice of $\alpha$ near 0 will produce efficiency close to that of the MLE but will lack in terms of robustness; on the other hand, the robustness increases with increase in $\alpha$. So, it is important to choose a proper $\alpha$ in any practical application of the proposed methodology. Based on our extensive simulations and empirical investigations, it has been observed that a choice of $\alpha\approx 0.3, 0.5$ provides the best trade-off despite the amount of contamination in data (which is often unknown); so it can serve as our empirical recommendation for the choice of $\alpha$ in any practical application of the proposed methodology. However, a data-driven algorithm for choosing an optimum $\alpha$ values for a given dataset could be useful in some complex contexts.
{\color{black} One such algorithm by minimizing and estimate of the asymptotic mean squared error has been explored by Ghosh/Basu:2015 in the context of linear regression models which can be directly adopted in our present context for estimating the ATE. As an alternative way, suggested by a reviewer, here we have briefly explored the use of bootstrap for the selection of optimal tuning parameter value $\alpha$. Such a bootstrap based approach has recently been used for tuning parameter selection in many problems; see, e.g., Ozkale/Altuner:2021. However, the standard application of bootstrap requires the assumption of independence between the observations, which is clearly not a valid assumption in our case with panel data. So, we resort to a more specialized approach to bootstrapping, called Moving Block Bootstrap Liu:1992. In this approach, for a sample of size $n$, we consider $n-b+1$ overlapping blocks of size $b$, each block containing $b$ consecutive data points. Then, instead of resampling from the observations as in usual bootstrap, we consider resampling the blocks instead, to preserve the correlation structure of the original observations to some extent. Finally, these blocks are combined together to form a bootstrap sample of size $n$. Hence, our empirical data-driven algorithm for selection of $\alpha$ using this approach of moving block bootstrap can be summarized as follows:
The rationale behind such a selection criteria is that we do not expect to have any kind of treatment effect at any of the pre-treatment time points. Hence, we would like to make the mean squared bootstrap estimates of ATE as close to 0 as possible.
We have briefly illustrated the applicability of this suggested procedure in case of our real data example in a later section. However, considering the contents and length of the present manuscript, more detailed exploration and verification of the performances of this suggested ways of selection of $\alpha$ under different set-ups are kept for a future work. }
Let us now further explore the robustness properties of the proposed MDPDE of the ATE theoretically via the classical influence function analysis. Let $G_{pre}(\boldsymbol{x},y_{1})$ and $G_{post}(\boldsymbol{x},y_{1})$ denote the joint distributions of $\boldsymbol{x}$ and $y_1$ in the pre-treatment and the post-treatment periods, respectively, where $\boldsymbol{x}=(1,\boldsymbol{\widetilde{y}})$. Then, the ATE estimator $\widehat{\Delta}_1^{(\alpha)}$ can be re-written in terms of their empirical estimator, namely $G_{n,pre}$ and $G_{n,pre}$, respectively, as follows:
where $\boldsymbol{U}_{\alpha}(.)$ denotes the functional for the MDPDE $\boldsymbol{\widehat{\beta}}_\alpha=(\widehat{\delta}_{1,\alpha},\boldsymbol{\widehat{\delta}_\alpha}).$ Therefore, the statistical functional corresponding to the Mean-MDPDE of the ATE would be defined as
To derive the influence function of this ATE functional, we consider the pre- and post-treatment contaminated distributions given by $$ G_{pre,\epsilon}=(1-\epsilon)G_{pre}+\epsilon \wedge(\boldsymbol{x}_t,y_{1t}) ~~~~\mbox{and}~~ G_{post,\epsilon}=(1-\epsilon)G_{post}+\epsilon \wedge(\boldsymbol{x}_t,y_{1t}), $$ where $\epsilon$ is the contamination proportion and $\wedge(.,.)$ denotes the distribution function for a degenerate distribution at the point of contamination $(\boldsymbol{x}_t,y_{1t})$. We will then derive the influence functions for the case of pre-treatment contamination and that of the post-treatment contamination in the following subsections.
In order to derive the influence function of the Mean-MDPDE ATE functional $T_{\alpha}^\ast\left(G_{pre},G_{post}\right)$, we need to substitute $G_{pre}$ by $G_{pre,\epsilon}$ in (ref) and differentiate it with respect to $\epsilon$. Evaluating the result at $\epsilon=0$, we get the desired influence function as
But, it can be easily shown that the influence function for the MDPDE functional $\boldsymbol{U}_{\alpha}$ at the normal model density is given by (see Basu/etc:2011,Ghosh/Basu:2013) $$ IF(\boldsymbol{x}_t,y_{1t});T,G_{pre})=(1+\alpha)^{3/2}E(\boldsymbol{X}_0'\boldsymbol{X}_0)^{-1}\boldsymbol{x}_t(y_{1t}-\beta'\boldsymbol{x}_t) e^{-\frac{\alpha(y_{1t}-\boldsymbol{\beta}'\boldsymbol{x}_t)^2}{2\sigma^2}}. $$ Therefore, the influence function for the Mean-MDPDE of the ATE under pre-treatment contamination turns out to be $$ IF_0((\boldsymbol{x}_t,y_{1t}), T_{\alpha}^\ast, (G_{pre},G_{post})) = -(1+\alpha)^{3/2}(y_{1t}-\boldsymbol{\beta}'\boldsymbol{x}_t)e^{-\frac{\alpha(y_{1t}-\beta'\boldsymbol{x}_t)^2}{2\sigma^2}} \int \boldsymbol{x}'E(\boldsymbol{X}_0'\boldsymbol{X}_0)^{-1}\boldsymbol{x}_t\hspace{1mm}dG_{post}(\boldsymbol{x}). $$ It is easy to see that the above influence function is bounded with respect to contamination ($y_{1t}$) in the response variable for all $\alpha>0$, justifying our empirical findings about the robustness of the Mean-MDPDE of the ATE. However, its boundedness, and hence the robustness, with respect to contamination in covariate space depends on the distribution of the covariates.
As a special case, the above influence function also provides the same for the classical HCW estimator of the ATE at $\alpha=0$, which turns out to be $$ IF_0((\boldsymbol{x}_t,y_{1t}), T_{0}^\ast, (G_{pre},G_{post})) = -(y_{1t}-\boldsymbol{\beta}'\boldsymbol{x}_t)\int \boldsymbol{x}'E(\boldsymbol{X}_0'\boldsymbol{X}_0)^{-1}\boldsymbol{x}_t\hspace{1mm}dG_{post}(\boldsymbol{x}). $$ This function is clearly unbounded even in $y_{1t}$ indicating the non-robust nature of the HCW estimator even under infinitesimal contamination.
\noindentAn Example:\\ For illustration, we plot an example of the influence function of the Mean-MDPDE of the ATE for different $\alpha>0$, and the same for the HCW estimator, in Figure (ref). Here, we have considered $\boldsymbol{x}=(1,y_2)'$, $\boldsymbol{x}_t=(t,t)$, $y_{1t}=t$ and $\delta_1=2$, $\delta=\beta_2=4$, $\sigma=2$. Clearly, the influence function of the classical HCW estimator is unbounded while we get a bounded influence function for the robust Mean-MDPDE at any tuning parameter $\alpha>0$. Further, the maximum value of these influence functions decreases with increasing values of $\alpha>0$ indicating the increase in the extent of robustness of the corresponding Mean-MDPDE of the ATE.
For contamination in the post-treatment distribution, we substitute $G_{post,\epsilon}$ in place of $G_{post}$ in (ref) and again differentiate with respect to $\epsilon$ at $\epsilon=0$. This leads to the influence function of the Mean-MDPDE of the ATE under post-treatment contamination as given by
Note that, this influence function is a linear function in the contamination points $\boldsymbol{x}_t$ and $y_{1t}$, and hence unbounded for any $\alpha\geq 0$. This further validates the results of our simulation study indicating that the Mean-MDPDE of the ATE is not suitable for the cases where we have contamination in the post treatment period.
\noindentAn Alternative Estimator for Post-Treatment Contamination:\\ Note that the main reason for the non-robustness of the Mean-MDPDE of the ATE under contamination in post-treatment period is the use of mean to get a summary measures from the individual ATE estimators $\widehat{\Delta}_{1t}^{(\alpha)}$ at the post-treatment period via Equation ((ref)), and mean is known to be a non-robust summary measure. So, in case of post-treatment contamination, we may get a robust summary estimate of the ATE by considering the median of $\widehat{\Delta}_{1t}^{(\alpha)}$s. This estimator, which we will refer to as the Median-MDPDE of the ATE, is formally defined as
To illustrate the claimed robustness of this Median-MDPDE of the ATE, we repeat the simulation study performed in Section (ref) to compare the empirical bias and MSE of $\widetilde{\Delta}_1^{(\alpha)}$. The results obtained for the post-treatment contamination are reported in Tables (ref) and (ref); results for the other cases are very similar to those obtained for the Mean-MDPDE of the ATE and hence they are not reported again for brevity. From these tables, it is evident that the Median-MDPDE of the ATE can successfully produce robust estimators in the presence of data contamination in post-treatment period in terms leading to significantly smaller MSE and bias compared to the Mean-MDPDE and all other existing ATE estimators considered here.
We will now illustrate the applicability of the proposed ATE estimator in a practical scenario of investigating the long term effects of the “2004 Indian Ocean Tsunami" Tsunami on the per capita GDP of the most severely affected countries. The tsunami of 2004 that originated in the Indian Ocean severely affected the surrounding countries causing damage to innumerable lives and properties. We want to estimate the effects of this natural calamity on the per capita GDP of the five severely hit countries which are Indonesia, India, Thailand, Maldives and Sri Lanka. For this purpose, we collect the data on per capita GDP (in US \$) from \url{https://www.macrotrends.net/countries/ranking/gdp-per capita}, over period 1981--2004 as the pre-treatment period and from 2005 to 2019 as the post-treatment period. As for the control group, it is natural to consider the same data for the countries which were not affected by the hazard. For our study, the countries in the control group are Portugal, Iceland, Israel, Netherlands, Brazil, USA, Japan, Ethiopia, Switzerland, Tunisia, Bangladesh, Pakistan, Philippines and Egypt, covering a wide range of economic conditions.
Based on some preliminary analyses, the logarithm of the per capita GDP of countries in the control group is taken as the covariate while the response is considered to be the logarithm of the per capita GDP of each of the affected (treated) countries. Then, we estimate the ATE of the Tsunami based on these data via the proposed Mean-MDPDE, Median-MDPDE and also using other existing ATE estimators; the resulting ATE estimates are reported in Table (ref). Note that, these results can be viewed as a long-term effects of Tsunami since they are computed summarizing the effects from 5 years in the post-treatment period. {\color{black} The values of optimum $\alpha$ obtained by the moving block bootstrap method, as described in Section (ref), and the corresponding Mean and Median-MDPDEs are also reported in the same Table (ref). }
{\color{black} Here we should point out that the HCW method required the data to be stationary which is not the case here with the logarithmic GDP values considered, but we have still kept this estimate among our results as a benchmark for comparison. On the contrary, our proposed estimators can indeed handle the non-stationary data and so are more appropriate for such applications. It can be seen from Table (ref) that our proposal gives a valid way to compute the average treatment effect which is comparable with other existing estimates in cases with no contamination and provide a better robust estimates in cases with possible data contamination.} That the proposed procedures give a better fit to the data with all moderate and optimum $\alpha$ values has been confirmed by looking at the corresponding coefficient of determination ($R^2$), which are not reported here for brevity.
To investigate further, we note that the estimated ATE based on the Mean-MDPDE (or the Median-MDPDE) with $\alpha>0$ are quite close to zero. So, we next test for the significance of these estimated ATEs by testing the hypothesis $H_0 : \Delta_1 = 0$ against $H_1 : \Delta_1 \neq 0$. We perform this test based on the proposed MDPDE-based Wald-type test statistics to get robust insights; the resulting $p$-values are plotted over the tuning parameter $\alpha\geq 0$ in Figure (ref). Note that, the p-value at $\alpha=0$ corresponds to the test of significance based on the existing HCW estimator. It is again evident that the null hypothesis of no significant effect is rejected for Indonesia and Sri Lanka at 5% level of significance (p-value for Thailand is also quite low as 0.108) while using the existing HCW estimator, but the proposed MDPDE based Wald-type tests, performed using the Mean-MDPDE of the ATE, give much more stable and high p-values for all the five countries. Based on these robust inference, we can then conclude that there is no significant average effect of tsunami on the per capita GDP of each of these countries over the period from 2005-2019.
It is important to note here that the above conclusion, although supported by the data, may appear to be non-intuitive at a first glance since such natural disasters are deemed to have negative economic consequences. However, the consequences of the 2004 tsunami were quite unexpected and so the subsequent international humanitarian responses towards the tsunami affected countries were swift and outstanding. Athukorala Athukorala:2012 wrote that it was the largest ever international aid commitment in terms of both the total volume of trade and aid per affected person. It also helped people of these developing countries not to suffer from diseases and starvation that usually follow such disasters. All these made it possible for these Tsunami-affected countries to recover promptly without having any long-term impact in their per capita GDP. In fact, it has also been argued that Heger/Neumayer:2019 such copious influx of international assistance triggered higher economic growth in the Aceh province of Indonesia than what would have happened in the absence of Tsunami. The same is also true for the other affected countries as supported by our empirical investigation based on the proposed robust procedures, {\color{black} but it definitely needs further economic investigation to confirm this empirical insights as there can indeed be many other hidden factors. But, this example clearly illustrates the usefulness of our proposed methodology in real data analyses which has been the main purpose of this paper.}
In real-life data, we often encounter unusual observations (outliers) that make analyses cumbersome. They often mislead the path of analysis if not treated properly. In order to circumvent their influence, we prefer robust methods that are not affected by the presence of data contamination. In the light of this need, we have developed a robust estimator of the ATE from panel data based on the popular approach of robust MDPDE. A rigorous study of the large sample properties of the new ATE estimator is carried out. Statistical tests of hypotheses based on the proposed ATE estimator are also developed. Theoretical study of the robustness of the proposed estimator is manifested through the influence function analysis. The real life performance of the new estimator is exemplified through the analysis of the effect of tsunami of 2004 on the per capita GDP of few severely hit countries where the traditional method failed to perform well in the post-treatment period but our proposed procedures have exhibited satisfactory results.
As mentioned before, we hope to study the data-driven algorithms for choosing an optimum $\alpha$ values for a given dataset in more details for complex contexts in our future research. We also hope to extend the proposed methodology for robust estimation of the ATE from other types of data in our sequel research works.
\noindentAcknowledgment:\\ This research is partially supported by an internal research grant from Indian Statistical Institute, India. Research of AG is also partially supported by an INSPIRE Faculty Research Grant from Department of Science and Technology, Government of India.