Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,789 characters · 14 sections · 92 citation commands
Dynamic covariate balancing: estimating treatment effects over time with potential local projections
\spacingset{1.7} Researchers collect a panel of $n$ independent observations observed over a finite number of $T$ periods in an observational study. The dataset encompasses time-varying covariates, outcomes, and time-varying treatments. The primary objective is to conduct inference on the average effect of exposure to different treatment histories, such as the effect of being treated for a certain number of periods.
We consider a setting where treatments change dynamically over time, and potential outcomes, covariates and treatment may depend on past histories. Two alternative procedures can be considered in this setting. First, researchers may consider explicitly modeling how treatment effects propagate over each period through time-varying covariates and intermediate outcomes. This approach is prone to large estimation error and misspecification in high-dimensions: it requires modeling outcomes and each time-varying covariate as a function of all past covariates, outcomes, and treatment assignments. A second approach is to use inverse-probability weighting estimators for estimation and inference tchetgen2012semiparametric, vansteelandt2014structural . However, classical semi-parametric estimators are prone to instability in the estimated propensity score. There are two main reasons. First of all, the propensity score defines the joint probability of the entire treatment history and can be close to zero for moderately long treatment histories. Additionally, the propensity score can be misspecified in observational studies.
This is a common problem both in social sciences and bio-statistics. For example, in a survey of all articles in 2021 top-5 economics journals, more than $20\%$ of studies with time-varying treatments exhibit treatment dynamics.\footnote{This is based on the authors' calculation. Top-5 economics journals are American Economic Review, Econometrica, Journal of Political Economy, Quarterly Journal of Economics, Review of Economic Studies.} On the other hand, typical approaches in economics and related disciplines often employ, Difference-in-Differences designs. If units can dynamically choose treatments in response to their previous outcomes (or treatments), this will lead to violations of the parallel trends assumption required by such designs ghanem2022selection, marx2022parallel. We, therefore, introduce an approach that is valid when treatment decisions at period $t$ can depend on the history of outcomes and treatments prior to $t$. The second challenge is that treatment dynamics are difficult to estimate. Individuals may select into treatment arbitrarily based on high-dimensional covariates, outcomes, and treatments, e.g., when maximizing future expected utilities heckman2007dynamic. This motivates a method that does not impose modeling assumptions on selection into treatment mechanisms (i.e., propensity score).
This paper studies the estimation and inference of the effects of treatment histories when potential outcomes (and covariates) depend on present and past treatments. Individuals dynamically select into treatment based on past (time-varying) covariates, outcomes, and treatments. There are no unobserved confounders after controlling for high-dimensional past characteristics ding2019bracketing. Researchers remain agnostic on the propensity score.
We leverage a model on the potential outcomes' conditional expectations as an (approximately) linear function of previous potential outcomes and (high-dimensional) covariates in each period. Our model is motivated by local projection frameworks jorda2005estimation, montiel2021local. Local projections impose a (linear) model on observed outcomes conditional on each period observables and do not require estimating how each time-varying covariate changes in response to treatments -- which would be prone to large estimation error in high dimensions. However, different from standard local projections, our model is imposed on expected potential instead of observed outcomes. This difference is important here because of treatments' serial correlation and selection into treatment based on past outcomes and covariates: a model on realized outcomes imposes restrictions on the distribution of the treatment assignments, whereas a potential outcome model does not. Building on the literature on marginal structural models robins2000marginal, we identify the parameters of interest by recursively projecting outcomes' conditional expectations over past histories, allowing for dynamic selection into treatment.
Our estimation method, Dynamic Covariate Balancing (DCB), estimates the parameters of the model by using recursive penalized projections through lasso hastie2015statistical. It then reweights observations to guarantee balance between treated and control units. Balancing covariates is intuitive and common in practice: in cross-sectional studies, treatment, and control units are comparable when the two groups have similar characteristics imai2014covariate, li2018balancing, hainmueller2012entropy. We generalize covariate balancing in the absence of dynamics of zubizarreta2015stable, athey2018approximate, ben2018augmented, hirshberg2017augmented to a dynamic setting. We show that balancing with potential local projections corresponds to constructing weights sequentially in time by first balancing treated and control units' covariates in the first period and then balancing histories in the next periods reweighted by the weights obtained in the previous period. The estimated balancing weights solve a sequence of quadratic programs to minimize the weights' variance.
Our estimation procedure guarantees a vanishing bias of order faster than $n^{-1/2}$ and a parametric rate of convergence of the estimated treatment effect in high-dimensional settings. In addition, the optimization problem over the set of balancing weights admits a feasible solution, with the true propensity score being one such solution (and without requiring knowledge of it). This result highlights the benefits of balancing over propensity score reweighting here: the proposed balancing weights have a smaller variance than inverse probability weights and -- by leveraging an (approximate) high-dimensional linear outcome model -- do not require the correct specification of the propensity score.\footnote{Typical methods in high dimensions require conditions on the product of the rates of estimators for the propensity score and coefficients of the linear model to be faster than $n^{-1/4}$, and also require consistent estimation of both the outcome model and propensity score model athey2017efficient. Compared to estimating the propensity score with a semi-parametric model, our guarantees do not depend on the estimation error of the propensity score (only require that the estimation error of the coefficients is $o(n^{-1/4})$), by leveraging the high dimensional linear outcome model. } This is an advantage especially in dynamic settings: the propensity score defines the joint probability that units are assigned to a given treatment history, and therefore inverse probability weights can exhibit large variance in finite sample (see e.g., Figure (ref)). Finally, we provide guarantees for inference. Relative to cross-sectional studies, our dynamic structure necessitates novel considerations for identification, balancing, and derivations, that require analyzing joint distributions of correlated residuals from sequential projections.
We illustrate our method in an empirical application using data from acemoglu2019democracy on studying the effects of democracy on economic growth. Here, the authors assume a dynamic selection model. Whereas effects are in magnitude and sign consistent with acemoglu2019democracy, we show that standard local projections and acemoglu2019democracy's linear regression lead to significantly smaller point estimates compared to our approach. We also show that (A)IPW methods lead to a more substantial imbalance (and bias) compared to DCB due to the instability of the propensity score in both high and low-dimensions.
The goal of this paper is to conduct inference on dynamic treatment effects, while being robust to the misspecification of the propensity score. To achieve this goal, we leverage a high dimensional linear model and derive the first dynamic balancing equations within the local projection model proposed in this paper.
In the econometrics and statistics literature, imbens2019panel propose balancing assuming no treatment dynamics, whereas here, treatment dynamics require different (and novel) balancing conditions. In the context of dynamics, different from imai2015robust, who estimate a single set of balancing weights over all possible combinations of time periods and covariates, here the number of moment conditions grows linearly with $T$ and not exponentially. Unlike zhou2018residual, who extend entropy balancing of hainmueller2012entropy to dynamic settings, and li2024toward who propose a single set of balancing by regressing each covariate on past information, we do not estimate one model for each covariate in the past (which can be prone to large estimation error in high dimensions). DCB explicitly characterizes the high-dimensional model's bias in a dynamic setting to avoid overly conservative moment conditions, while kallus2018optimal design conservative balancing conditions for the worst-case bias. Different from yiu2018covariate, we do not require estimating the propensity score. Our insight with respect to all these references (in low and high dimensions) is that with a linear model and sequential weights, balancing reduces to few and novel dynamic restrictions. This insight is even more relevant with high-dimensional covariates, which none of these references study with dynamics.
Compared to cross-sectional studies, we generalize balancing in athey2018approximate, ben2018augmented, and consider an arbitrary class of weights. Therefore, our residual balancing procedure does not reduce to linear estimators as in settings with linear balancing weights bruns2023augmented, and our analysis differs from cross-sectional studies with low dimensions in wang2020minimal.
More broadly, this paper connects to the literature on DiD, local projections, and dynamic treatments. Different from the literature on DiD rambachan2023more, de2022difference, callaway2019difference, abraham2018estimating, athey2022design, caetano2022difference or subsequent work on potential projections with DiD dube2023local, here we allow for dynamic treatment regimes. This literature imposes the parallel trends assumption, violated with dynamic treatments marx2022parallel, ghanem2022selection. Different from the time-series literature montiel2021local, stock2018identification, rambachan2019nonparametric, this paper uses information from panel data and allows for arbitrary dependence of outcomes, covariates, and treatment assignments over time.
References in bio-statistics include robins2000marginal, hernan2001marginal, boruvka2018assessing, blackwell2013framework, bang2005doubly vansteelandt2014structural. bojinov2020panel study IPW estimators from a design-based perspective. Doubly robust estimators for dynamic treatments have been studied by nie2021learning, zhang2013robust, jiang2015doubly, tchetgen2012semiparametric, babino2019multiple. Here, we focus on studying the effect of a given treatment path, as in robins2000marginal or blackwell2013framework, different from and complementary to studying optimal policies (e.g., murphy2003optimal, nie2021learning).
Specifically, studies with high-dimensional panels require correct specification of the propensity score lewis2020double, zhu2017high, shi2018high, bodoryevaluating, belloni2016inference, chernozhukov2017orthogonal, or impose homogeneous treatment effects high_dim_IRF, kock2015inference. (lewis2020double also illustrate bounds on misspecification). More generally, prior works that formally study properties of dynamic doubly-robust methods in high dimensions require product of rates conditions for the estimated propensity score and conditional mean function, and consistent estimation of both; see for example follow up work by bradic2021high who provide tight rates of convergence of dynamic AIPW. Different from above, our framework does not require consistent estimation of the propensity score.
Finally, in both works subsequent to the first version of this paper, chernozhukov2022automatic generalize the use of riesz representers with arbitrary non-linear outcome models, and zhang2021dynamic study doubly robustness to model misspecification through moment restrictions. Different from these references, here we do not require conditions on the balancing weights motivated by our goal of allowing for inference with a possibly completely misspecified propensity score function, whereas chernozhukov2022automatic and zhang2021dynamic require functional form restrictions on the balancing weights to obtain a product of rates conditions. Our focus on the high-dimensional linear model (which we view as a linear approximation to conditional expectations in high dimensions) is motivated by its large use in applications.
We start with the analysis of two time periods, deferring multiple periods to Section (ref). We observe a panel with $n$ $i.i.d.$ copies of $ \Big( X_{i,1}, D_{i, 1}, Y_{i, 1}, X_{i, 2}, D_{i, 2}, Y_{i, 2} \Big)$, each distributed according to $\mathcal{P}$. Here $ D_{i,1}, D_{i,2} \in \{0,1\}$ denote binary treatments at time $t = 1,t = 2$, respectively, $X_{i,t}, Y_{i,t}$ denote covariates and the outcome at time $t$. We allow for any nonstationarity and dependencies that may occur over time within each unit. When indices are not specified, such as in \(D_t\), this refers to the collective observations for all $n$ units.
We consider potential outcomes that are functions of the entire treatment history with $Y_{i,2}(d_1, d_2)$ denoting the potential outcome at time $t = 2$, under treatment $d_1$ in the first and $d_2$ in the second period. Our goal is to conduct inference on the estimand(s) $$
$$ for given treatment histories $(d_1, d_2), (d_1', d_2')$. For example, researchers may be interested in estimating $ \mathrm{ATE}((1,1), (0,0)), $ which denotes the \textit{total} effect of treating an individual for two consecutive periods \citep{athey2022design}; or the \textit{direct} effect $\mathrm{ATE}((1, 0), (0,0))$. Figure (ref) shows that the overall treatment effects capture the direct effect of the treatment on the outcomes and the indirect effect. For longer histories, one could also consider weighted combinations of relevant treatment effects, omitted for brevity (see Section (ref)).
Treatment histories can impact both outcomes and covariates at intermediate stages. Let \(Y_{i,1}(d_1, d_2)\) represent the intermediate potential outcome and \(X_{i,2}(d_1, d_2)\) represent the potential covariates following a sequence of \(d_1\) then \(d_2\). Here, \(X_{i,1}\) refers to the baseline covariates.
Assumption (ref) is a no-anticipation restriction: (i) intermediate potential outcomes only depend on past but not future treatments; (ii) the treatment status at $t = 2$ has no contemporaneous effect on covariates.
Assumption (ref) allows for anticipatory effects governed by expectations (e.g., individuals may choose treatments based on expected future utilities), but not on the future treatment realizations athey2022design.
In the rest of our discussion, we index potential outcomes and covariates by past treatment history under Assumption (ref). We define $ H_{i,2} = \Big[D_{i,1}, X_{i,1}, X_{i,2}, Y_{i,1}\Big], $ the vector of past treatment assignments, covariates, and outcomes in the previous period. We refer to $ H_{i,2}(d_1) = \Big[d_1, X_{i,1}, X_{i,2}(d_1), Y_{i,1}(d_1)\Big] $ as the potential history under treatment status $d_1$ in the first period. Here, $H_{i,2}$ can include interaction terms, omitted for brevity.
Sequential ignorability states that treatment in the first period is unconfounded conditional on baseline covariates, and the treatment in the second period is unconfounded conditional on all observable characteristics at $t= 2$. It assumes no unobserved factors after controlling for high dimensional observable characteristics and arbitrary past information. Note that we could also state (A), conditioning on $D_{i,1} = d_1$ and potential history $H_{i,1}(d_1)$.
In Example (ref), Assumption (ref) holds if $ D_{i,2} \perp \varepsilon_{i,2} \Big| D_{1, i}, X_{i,1}, X_{i,2}, Y_{i,1}, \quad D_{i, 1} \perp (\varepsilon_{i,1}, \varepsilon_{i,2}) \Big| X_{i,1}. $
Following in spirit, jorda2005estimation, we approximate the expectation of potential outcomes as linear functions of (high-dimensional) past characteristics. Different from jorda2005estimation, linearity is imposed on expected potential instead of realized outcomes. Modeling potential outcomes directly avoids functional form restrictions on the treatment assignment mechanism.
Assumption (ref) allows for heterogeneity in $(d_1, d_2)$, and the dimensions $p_1, p_2$ can grow with $n$ (because of additional covariates and/or covariates transformations). As for MSMs robins2000marginal, Assumption (ref) (i) does not require estimating a structural model for each time-varying-covariate, that would be prone to large estimation error in high dimensions; and (ii) it is agnostic on the treatment assignment mechanism because the model is imposed on potential outcomes. Coefficients can vary with time in the model.
The proof is in Appendix (ref). Lemma (ref) builds on results in the literature on marginal structural models robins2000marginal, bang2005doubly, tran2019double, kallus2020double, where, here, we make a connection between marginal structural models and local projections in economics as a contribution of independent interest. Lemma (ref) motivates a recursive estimation strategy discussed in Section (ref).
This section studies estimation. We defer to Section (ref) a complete guide for practice, including discussion about the model, tuning parameters, and complexity. Appendix (ref) presents a more detailed description of each step to construct the estimator.
Consider a two periods setting first, with $T = 2$. The estimator for $\mu(d_1,d_2)$ (and symmetrically for $\mu(d_1', d_2')$) proceeds in the following steps:
We formalize this discussion below (Appendix (ref) contains a formal proof).
We now describe in details the procedure with finite $T$ periods. Let $d_{1:T} = (d_1, \cdots, d_T)$,
This estimand denotes the difference in potential outcomes for two treatment histories $d_{1:T}, d_{1:T}'$. We denote
the vector containing information from time one to time $t$, after excluding the treatment assigned in the present period $D_t$. Interaction components may also be considered, omitted here for brevity. We let the potential history be $ H_{i,t}(d_{1:(t-1)}) = \Big[d_{1:(t-1)}, X_{i,1:t}(d_{1:(t-1)}), Y_{i,1:(t-1)}(d_{1:(t-1)}) \Big]. $
Assumption (ref) generalizes Assumptions (ref)-(ref) from the two-period setting. Identification follows similarly to Lemma (ref). For given weights $\hat{\gamma}_{1:T}$, and coefficients $\hat{\beta}^{(1:T)}$, we estimate
Coefficients are estimated recursively as in the two periods setting (see Algorithm 2).
The proof is in Appendix (ref). Lemma (ref) decomposes the estimation error into three components. First, ($I_1$), depends on the estimation error of the coefficient and on balancing properties of the weights. ($I_1$) suggests imposing balancing conditions on \\ $ \Big| \Big| \hat{\gamma}_t(d_{1:T}) H_t - \hat{\gamma}_{t-1}(d_{1:T}) H_t \Big| \Big|_{\infty} $ each period. The components characterizing the estimation error are $ (I_2)=\hat{\gamma}_T(d_{1:T})^\top \varepsilon_T$, and ($I_3$), which are mean zero as described in Appendix Lemma (ref). Note that $\bar{X}_1 \beta_{d_{1:T}}^{(1)}$ converges to $\mu_T(d_{1:T})$ under standard $\sqrt{n}$-asymptotics.
Next, we study the theoretical properties of the estimator in finite $T$ periods. We consider a high dimensional regime where the dimension covariates in each period $p_1, \cdots, p_T$ can grow to infinity, as long as $\log (\max_t p_t n)/n^{1/4} \to 0$. We impose the following conditions.
Condition (i) is the overlap condition, standard in the causal inference literature. The overlap condition is sufficient (but not necessary, see the discussion of athey2018approximate in cross-sectional settings) to show existence of a feasible solution of Algorithm 1; see Remark (ref). Condition (ii) is a tail restriction. Assumption (ref) can be relaxed by assuming that the product of the inverse probability weights times the covariates is sub-exponential at the expense of more tedious derivations.
The proof is in Appendix (ref). In particular, the reader may refer to Appendix Lemma (ref) for additional details on the construction of the feasible weights. The algorithm thus finds weights that minimize the small sample variance, with the IPW weights $w_{i,t}^*$ (reweighted by the solution in the previous iteration $\hat{\gamma}_{i,t}$) being one possible solution. Minimizing the $l_2$ norms of the weights is a natural objective when the goal is to minimize the variance of the ATE estimator: following Theorem (ref) below, under homoskedasticity of the residuals from each projection, the variance is proportional to a weighted sum of $||\hat{\gamma}_t||^2$. However, homoskedasticity is not necessary for our results.
Corollary (ref) provides a desiderable stability property. It shows that the $l_2$ norm of the weights is upper bounded by the (stabilized) IPW weights reweighted by the solution in the previous step. From basic concentration inequalities, for $n$ sufficiently large, $\mathbb{E}[\hat{\gamma}_{i,t}^{*2}(\hat{\gamma}_{t-1})] \approx \mathbb{E}\Big[\frac{ \hat{\gamma}_{i,t-1}^2 }{P(D_{i,t} = d_t | H_{t})}\Big]$. In addition, the last part of the corollary shows that the weights' norm is controlled over each period by the norm in the previous period up to a finite multiplicative constant.
Assumption (ref) imposes the consistency in estimating the outcome models. Condition (i) is attained for many high-dimensional estimators, such as the lasso method, under sparsity and restricted eigenvalues restrictions; see, e.g., buhlmann2011statistics. An example and derivation for condition (i) for Lasso under sparsity is included in Example (ref) (Appendix (ref)). As we discuss in Appendix Remark (ref), the restricted eigenvalue condition is imposed for all $H_t, t \ge 1$; references that study this condition from different angles include deshpande2023online, Section C.4.
Theorem (ref) guarantees a parametric convergence rate with high-dimensional covariates.
Inference on ATE follows as a direct corollary for two histories $d_{1:T}, d_{1:T}'$ with $d_1 \neq d_1'$ (see Theorem (ref)), as described in Appendix (ref). Also, for inference conditional on baseline covariates in the first period $X_1$, the relevant variance is $\hat{V}_T(d_{1:T}) - \frac{1}{n} \sum_i (\bar{X}_1 \hat{\beta}_{d_{1:T}}^{(1)} - X_{i,1} \hat{\beta}_{d_{1:T}}^{(1)})^2$ (since we condition on $X_1$) and for the corresponding ATE is the sum of these two variances.
The complete Algorithm 1 is implemented off-the-shelf in the R-package {\tt DynBalancing}.
It requires researchers to specify four main parameters: the length $h$ of the treatment history considered (i.e., carry-over effects), two treatment histories of length $h$, $d_{(T-h):T}, d_{(T-h):T}'$ to compare, the model used to estimate the coefficients ({\tt linear} or {\tt fully interacted}) as described in Algorithm 2, and whether to consider a {\tt pooled} regression.
Choosing the length of the treatment history with long panel With short panels, selecting the length of the treatment history $h = T$ is natural. With long panels, this may reduce the effective sample size or be infeasible (as the effective sample becomes “thinner"). This is because, as for IPW, the weights at time $t$ can be non zero only for those units observed over a given treatment path up to time $t$. Therefore, we recommend selecting a treatment history $h$ shorter than the number of periods $T$ (i.e., $h < T$), and estimate causal effects of the form
for given treatment histories $d_{(T-h):T}, d_{(T-h):T}'$. Equation (ref) estimates the effect of exposing an individual to two different histories over the last $h$ periods and average over previous assignments. Our analysis and estimation follow similarly to Algorithm 1, with the difference that we construct balancing weights starting from period $T - h$ and proceed sequentially until time $T$ (observable characteristics before time $T - h$ can be used as additional controls). As in imai2018matching, the focus on Equation (ref) makes our procedure robust to long panels.
As a rule of thumb, as we illustrate in our application, we recommend report results for different choices of $h$ (say $h \in \{1, \cdots, 10\}$ in a long panel); our package reports and plots the estimated effects along-side standard errors which can help disentangle the trade-offs between identification of long-run effects against precision. In addition, it is useful to report $1/(n ||\hat{\gamma}_t||^2)$ as a measure of effective sample size at time $t$ for different values of $t$, which can help guide the choice of $h$ (larger choices of $1/(n ||\hat{\gamma}_t||^2)$ indicates more accurate treatment effects estimates). \\ Choice of the model specification ({\tt linear} or {\tt fully interacted}) The estimation error $||\hat{\beta}_{d_{1:T}}^{(t)} - \beta_{d_{1:T}}^{(t)}||_1$ depends on modeling assumptions. For the {\tt fully interacted} model, $||\hat{\beta}_{d_{1:T}}^{(t)} - \beta_{d_{1:T}}^{(t)}||_1$ scales exponentially with $T$ as it considers all possible interactions with the treatment assignments $d_1, \cdots, d_T$. The {\tt linear} model avoids that the effective sample size shrinks exponentially in $T$ but imposes homogeneity restrictions of treatment effects as for example in acemoglu2019democracy, by modeling treatment effects as additive and linear. See Algorithm 2 for more details. In addition, when {\tt pooled} is true, we consider a regression $$
$$ where $\tau_t$ denotes fixed effects, pooling together effects estimated in different periods. We then cluster standard errors at the individual level to allow for correlation over time.
Next, we collect results from numerical experiments. We estimate $ \mathbb{E}\Big[Y_{i,T}(1, \cdots, 1) - Y_{i,T}(0, \cdots,0)\Big], T \in \{2,3\}. $ We let the baseline covariates $X_{i,1}$ be drawn from as i.i.d. $\mathcal{N}(0, \Sigma)$ with $\Sigma^{(i,j)} = 0.5^{|i-j|}$. Covariates in the subsequent period are generated according to an auto-regressive model $ \{X_{i,t}\}_j =0.5 \{X_{i,t-1}\}_j + \mathcal{N}(0, 1), j=1,\cdots,p_t. $ Treatments are drawn from a logistic model that depends on all previous treatments and past covariates: $ D_{i,t} \sim \mbox{Bern}\Big((1 + e^{\iota_{i,t}})^{-1}\Big) $ with
and $\xi_{i,t} \sim \mathcal{N}(0,1)$, for $t \in \{1, 2,3\}$. Here, $\eta, \delta$ controls the association between covariates and treatment assignments. We consider values of $\eta \in \{0.1, 0.3, 0.5\}$, $\delta_1 = 0.5, \delta_2 = 0.25$. We let $\phi \propto 1/j$, with $\|\phi \|_2^2 = 1$, similarly to balancing conditions presented in athey2018approximate. The larger $\eta$ corresponds to weaker overlap (see Table (ref) in the Appendix).
We generate the outcome as $ Y_{i,t}(d_{1:t}) = \sum_{s = 1}^t \Big(X_{i,s} \beta + \lambda_{s, t} Y_{i,s-1} + \tau d_s\Big) + \varepsilon_{i,t}(d_{1:t}), \quad t=1,2,3, $ where elements of $\varepsilon_{i,t}(d_{1:t})$ are i.i.d. $ \mathcal{N}(0,1)$ and $\lambda_{1,2} = 1, \lambda_{1,3}, \lambda_{2,3} = 0.5$. We consider three different settings: Sparse with $\beta^{(j)} \propto 1\{j \le 10\}$, Moderate with moderately sparse $\beta^{(j)} \propto 1/j^2$ and the Harmonic setting with $\beta^{(j)} \propto 1/j$. We set $\| \beta \|_2 =1, \tau = 1$.
We consider the following competing methodologies:
For Dynamic Covariate Balancing, DCB, the choice of tuning parameters is data adaptive, and it uses a grid-search method discussed in Appendix (ref) and Remark (ref). We estimate coefficients as in Algorithm 1 for DCB and (a)IPW, with a linear model in treatment assignments. Estimation of the penalty for the lasso methods is performed via cross-validation.
We consider $\mathrm{dim}(\beta) = \mathrm{dim}(\phi) = 100$ and set the sample size to be $n = 400$. We set $p_1=101, p_2=203, p_3=305$ as number of covariates in each period.
In Table (ref) we collect results for the average mean squared error for estimating the average treatment effect in two and three periods. Throughout all simulations, the proposed method significantly outperforms any other competitor, with one single exception for $T = 2$, good overlap and harmonic design. It also outperforms using known propensity score, consistently with our findings in Theorem (ref), where we show that the propensity score is a feasible solution of DCB weights (and in the absence of knowledge of the propensity score).
Finally, in Appendix (ref) we consider more extensive simulation studies with a longer time horizon, a misspecified (non-linear) model, low and high dimensional settings among additional simulation designs.
\definecolor{glaucous}{rgb}{0.38,0.51,0.71}
In this section, we present an empirical application for studying the effect of democracy on GDP growth using data from acemoglu2019democracy. acemoglu2019democracy studied dynamic treatment effects of democracy under GDP growth under sequential ignorability acemoglu2019democracy. Figure (ref) illustrates the dynamics of treatments. Many units switch treatment over time, violating standard event studies designs.
The data (available at \url{https://www.journals.uchicago.edu/doi/suppl/10.1086/700936}) consist of a collection of countries observed between $1960$ and $2010$. We consider observations starting from $1989$. After removing missing values, we run regressions with 141 countries. The outcome is the log-GDP in the country $i$ in period $t$ as in acemoglu2019democracy. We use the same treatment specification as in acemoglu2019democracy, which is binary. We study the effect of exposing countries at time $t$ to democracy for in $s$ years before (and including) $t$ versus not exposing them to democracy for the previous $s$ years. Namely, the estimand is the $s$-long run effect of democracy, after averaging over past assignments. We let $s \in \{1, \cdots, 20\}$ to study the impact from one to twenty years of democracy.
For each country, we condition on lag outcomes in the past four years as in the preferred specification of acemoglu2019democracy, and past four treatments. We consider a pooled regression and two alternative specifications. The first is parsimonious and includes dummies for different regions (continents) and different intercepts for different periods. The second one controls for the past four outcomes, past four treatments, for the geographical region, and colonial history as in acemoglu2019democracy. Coefficients are estimated as in Algorithm 2 with {\tt model} $=$ {\tt linear}.
\definecolor{aero}{rgb}{0.49,0.73,0.91} \definecolor{airsuperiorityblue}{rgb}{0.45,0.63,0.76} \definecolor{babyblueeyes}{rgb}{0.63,0.79,0.95} \definecolor{beaublue}{rgb}{0.74,0.83,0.9} \definecolor{glaucous}{rgb}{0.38,0.51,0.71}
Figure (ref) collects our results. Democracy has a statistically insignificant effect over the first few years and a statistically significant positive impact on long-run GDP growth after three years. Point estimates are in sign and magnitude consistent with what found by acemoglu2019democracy, and results are robust across the two specifications for DCB.
We compare our method to (i) the linear estimator reported by acemoglu2019democracy (Table 2, Column 3), where dynamic effects are estimated by propagating the effect over past outcomes at each period (we consider two specifications, with and without unit fixed effects -- both report similar results); (ii) the simple local projection, that projects the outcome on the treatment and the past outcome $s$ periods before, with and without country fixed effects, time fixed effects and controlling for lagged outcome at time $s$.
The simple local projection approach reports small point estimates compared to other methods. This result is consistent with our theoretical discussion: local projections average over the distribution of future assignments. Therefore, the causal effects estimated by the local projection differ from the target long-run effect, which instead fixes future treatment assignments. The effect estimated as in acemoglu2019democracy is larger than the local projection when including country fixed effects, but significantly smaller than the effect estimated through DCB. Therefore, the specification in acemoglu2019democracy may capture some but not all the long-run effects. After controlling for imbalance with DCB, average treatment effects are twice as large. The results from acemoglu2019democracy with and without unit fixed effects report almost identical results.
To investigate differences with (A)IPW methods, the right panel in Figure (ref) presents comparisons in terms of the imbalance over the lagged outcome at time $t - 1$ when using balancing or inverse probability weights. As acemoglu2019democracy note, the lags outcome may capture most variation in treatment. Therefore, an imbalance in lagged GDP may suggest the presence of bias. We report the relative improvement in absolute imbalance (average across the potential outcomes under treatment and control) and observe substantial gain over using inverse probability weights. Such gains illustrate the advantage of balancing in small sample.
Figure (ref) complements Figure (ref) showing instability of inverse probability weights, and Figure (ref) in the Appendix show that DCB weights present less dispersion than IPW weights. As a helpful diagnostic, in Table (ref) we report $1/(||\hat{\gamma}_t||^2)$ for the DCB method as well as for IPW weights, where $\hat{\gamma}_t$ is replaced by the AIPW weight. This measure is indicative of the level of precision and effective sample size (a larger number indicates better precision). We find substantial improvements of DCB over IPW especially for weights estimated for longer time horizons.