Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
39,538 characters · 13 sections · 5 citation commands
Reducing bias in difference-in-differences models using entropy balancing
Assessing the impact of health policies frequently requires causal inference using observational research designs because randomized controlled trials are infeasible. When both pre- and post-intervention outcomes for the intervention and comparison groups are available, difference-in-differences (DD) is a widely used research design for causal inference.\supercite{wherry2016early,hanchate2015massjoint,osborne2015nsqip,lasser2014massreadmit,smulowitz2011massED,mcwilliams2016aco,mcwilliams2013aco} DD uses the difference between intervention and comparison groups in pre-intervention outcomes to control for permanent differences across groups in unobservable factors that affect outcomes; changes in the difference in group means after the intervention are attributed to the intervention.\supercite{wooldridge2002econometric} DD delivers unbiased estimates of causal effects under the assumption that the intervention and comparison groups would experience identical changes in post-intervention outcomes over time in the absence of the intervention, an assumption commonly referred to as parallel trends.\supercite{angrist2008mostly}
Although it is possible to estimate intervention effects using DD with data on group-level mean outcomes, it is common for researchers to carry out a DD study by estimating a regression model on individual-level microdata, sometimes referred to as the microlevel DD estimator.\supercite{daw2018matching} Among other advantages, the use of microdata allows researchers to control for individual-level covariates, making the DD design more robust to observable changes in group composition. In the microlevel DD, the parallel trends assumption on unconditional group means can be replaced with a weaker assumption that post-intervention trends in group outcomes conditional on included covariates would be parallel in the absence of the intervention.\supercite{angrist2008mostly, stuart2014using} Regardless of the exact form, the parallel trends assumption is inherently untestable since the outcomes for the intervention group absent the intervention can never be observed.
Recently, researchers have combined matching with microlevel DD to further improve balance between groups on baseline observables, potentially including matching on pre-intervention outcomes.\supercite{stuart2014using} Whether matching on pre-intervention trends or other baseline observables actually reduces the bias of DD estimators is subject to debate. Prior simulation studies have shown that matching combined with a DD analysis can lead to bias reduction when compared to the usual DD analysis.\supercite{lindner2019difference,ryan2015we} Others, meanwhile, have argued that matching on pre-intervention outcome levels is insufficient to eliminate bias.\supercite{chabe2015analysis} In a recent paper, daw2018matching\,\supercite{daw2018matching} point out that naively matching time-varying covariates or pre-intervention outcome levels can actually increase bias. daw2018matching\,\supercite{daw2018matching} also considered matching on pre-intervention trends and found that matching on pre-period trends failed to eliminate bias. They recommend that researchers should not use a DD study design when pre-intervention trends between the intervention and comparison groups are not parallel. In related work, arkhangelsky2019synthetic\,\supercite{arkhangelsky2019synthetic} propose a synthetic DD approach that weights comparison subjects based on pre-intervention outcome levels.
In this paper, we reexamine the questions of whether, and how, to combine DD with matching, weighting, or other approaches designed to balance pre-intervention outcome trends and characteristics. We consider settings where researchers have microdata covering multiple pre-intervention time periods, making it possible to reweight the comparison group based on information about individual-level trends in pre-intervention outcomes. We discuss assumptions under which weighting can remove bias relative to unweighted DD and describe a set of sufficient conditions for such estimators to yield unbiased estimates of causal effects. As an alternative to matching, we show how to use a weighting method known as entropy balancing to obtain balance on pre-intervention outcome trends.\supercite{hainmueller2012entropy} Using a simulation analysis, we demonstrate that it is possible to obtain unbiased DD estimates of causal effects even when the parallel trends assumption is not satisfied by the unweighted group means.
We present an empirical example drawn from a recently completed policy evaluation to illustrate the use of entropy balancing in a DD study. Under the Value-Based Insurance Design (VBID) Model Test, insurers offering Medicare Advantage (MA) coverage (Medicare Advantage Organizations, or MAOs) in certain states were allowed to modify benefit design based on a patient's health status, thereby encouraging patients with specified chronic conditions to increase utilization of high-value care and avoid costly and harmful complications. Using MA encounter data from 2014 to 2017, we estimate the impact of MA VBID on the number of primary care visits and the number of specialty care visits among VBID-eligible beneficiaries.
We define the causal effect of interest using potential outcomes in continuous time.\supercite{rubin2005causal} Let $A_i$ denote a binary indicator that the $i$-th individual was subject to the intervention. Further, assume that outcome data is available for all times $t \geq 0$, and the intervention is applied at time $t=t_e$, with $t_e > 0 $. We focus on instantaneous effects at a fixed time point $t \geq t_e$, but our analysis also applies to time-averaged effects defined over a range of post-intervention time points.
Let $Y_{i}(1,t)$ and $Y_{i}(0,t)$ denote the potential outcomes that would be observed with and without the intervention, respectively, for the $i$-th individual at time $t$. We observe only one potential outcome at each time point $t$, which is denoted by the observed outcome $Y_{i}(t) = A_i Y_{i}(1,t) + (1-A_i) Y_{i}(0,t)$.
For individual $i$, the instantaneous effect at time $t$ is defined as the difference in their potential outcomes,
Averaging these individual-level instantaneous effects over the population that was subject to the intervention defines the estimand of interest, the average treatment effect on the treated (ATT) at time $t$,
Microlevel DD can be formulated in many different ways, including through differences in groups means. For presentation, we define DD models as regression models. A microlevel DD model that specifies common time effects as a polynomial of order $P$ can be written as follows:
where $E_t$ is an indicator for the post-intervention time period ($t \geq t_e$), $\alpha$ is a constant, and $\varepsilon_{it}$ is an error term. Alternatively, a microlevel DD model with nonparametric time effects can be written as follows:
where $\mu_t$ denotes fixed effects for each time period and $\varepsilon_{it}$ is an error term.
In our simulation analysis, we consider DD models that control for time linearly (Equation (ref) with $P=1$), quadratically (Equation (ref) with $P=2$), or nonparametrically (Equation (ref)). However, the theoretical arguments in the remainder of this section are not tied to any one specification of DD.
The ATT is a function of unobservable potential outcomes, and is not identified from observed data without additional assumptions. The key identifying assumption for DD is the parallel trends assumption. To formalize this assumption in the context of weighted DD, we define $\delta_{i}(a,t) \equiv \frac{\partial}{\partial t} Y_i(a,t)$ as the potential outcome trend for individual $i$ under intervention assignment $a$ at time $t$. $\delta_{i}(0,t)$ is the potential outcome trend absent the intervention, and $\delta_{i}(1,t)$ is the potential outcome trend under intervention. The parallel trends assumption for weighted DD is as follows:
Assumption (ref) states that it is possible to find a sub-population of the comparison group that, after weighting, exhibits identical post-intervention potential outcome trends to those in the intervention group absent the intervention. In the Appendix, we prove that Assumption (ref) is sufficient to identify the average treatment effect on the treated using DD.
The parallel trends assumption is untestable because it depends on unobserved potential outcomes. It also provides no indication of how to find $w_i$. To guide the derivation of $w_i$ in practice, we introduce two alternative assumptions involving pre-intervention outcome trends that, together, imply Assumption (ref).
Assumption (ref) states that, among the comparison group, it is possible to find a weighted sub-population that exhibits similar pre-intervention outcome trends to the intervention group. Unlike Assumption 1, this is a testable assumption, as it relates to observed outcomes only.
In order for Assumption (ref) to have any implications for the validity of DD, it is necessary to make a further assumption about the relationship between pre- and post-intervention potential outcomes.
Assumption (ref) states that if the groups have parallel potential outcome trends in the pre-intervention period, then they also have parallel post-intervention potential outcome trends absent the intervention. Assumption (ref) allows arbitrary changes in outcome trends after the intervention, provided that such changes would affect both groups absent the intervention, i.e., it allows for common shocks. Assumption (ref) formalizes the intuition that motivates the standard practice of testing for parallel pre-intervention trends in DD studies with multiple pre-intervention time periods.
Together, Assumptions (ref) and (ref) imply Assumption (ref). Assumptions (ref) and (ref) are thus jointly sufficient for DD to identify the ATT (see Appendix). Furthermore, the sufficiency of Assumptions (ref) and (ref) for the DD to estimate the ATT provides a path forward for deriving the weights $w_i$: $w_i$ should be chosen to balance pre-intervention outcome trends between the intervention and comparison groups. This can be done using existing approaches designed to balance pre-intervention characteristics, including matching,\supercite{daw2018matching} propensity score methods,\supercite{stuart2014using,linden2011applying} or entropy balancing. Below we discuss how to incorporate entropy balancing into weighted DD estimation.
We describe the application of Hainmueller's\supercite{hainmueller2012entropy} entropy balancing approach to calculating weights for DD that balance pre-intervention outcome trends, thereby satisfying Assumption (ref). Methodological details are left to the Appendix. Entropy balancing selects weights that satisfy a series of balancing constraints based on the distribution of the observables in intervention and comparison groups. Although previous studies have combined entropy balancing with DD, these applications have focused on balancing the levels of pre-intervention covariates rather than pre-intervention outcome trends.\supercite{marcus2013effect,parish2018using} The application of entropy balancing to pre-intervention trends in DD has not previously been analyzed in detail.
For DD, the most important variables to include in the balancing constraints are estimates of the pre-intervention outcome trends. We consider two approaches to estimating individual-level pre-intervention outcome trends for inclusion in the balancing constraints (see Appendix for details). The first strategy is to model the trends parametrically, for instance using linear regression to estimate a pre-intervention outcome trend for each individual. We expect this method will perform well when the individual-level trends are noisy, as the parametric trends will average out the noise over time. However, misspecification of the parametric trends used to define entropy-balancing weights may result in violation of the Assumption (ref) or (ref) since the misspecified trends do not represent the true trends.
An alternative is to nonparametrically estimate the slope between each time point by taking first differences of the pre-intervention observed outcomes. We expect this method will perform well when observed outcome trends reflect heterogeneity across individuals, rather than transitory within-individual variability (or noise) that weakens the relationship between pre-intervention observed outcome trends and post-intervention potential outcome trends.
Regardless of how the pre-intervention outcome trends are estimated, Assumption (ref) relates to the true pre-intervention outcome trends while, in practice, we attempt to balance estimated pre-intervention outcome trends. If individual trends in outcomes reflect noise---perhaps due to measurement error or other transitory error components---then estimated outcome trends may be less reliable (in the sense of Steiner et al.\supercite{steiner2011importance}) and we expect that our ability to remove bias will suffer.
For entropy balancing to allow DD estimation of the ATT, one more assumption is required\supercite{hainmueller2012entropy}: the support of the intervention group's distribution of variables used in the balancing constraints must lie within the support for the comparison group, which we refer to as overlap. Overlap assumptions are standard in the literature on matching and related causal inference methods.\supercite{stuart2010matching} In our context, this assumption applies to the distribution of pre-intervention trends:
Overlap may be violated if there are individuals in the intervention group with pre-intervention outcome trends that differ radically from the comparison group. In such cases, it may not be possible to find a set of weights that allows estimation of the ATT. In this case, it may still be possible to estimate a local average treatment effect defined on the subpopulation of the intervention group that satisfies overlap.
We demonstrate the use of weighted DD estimation using simulated data in order to evaluate the performance of alternative estimators and to highlight the assumptions necessary to reduce bias by balancing pre-intervention outcome trends. Technical details of these simulations are presented in the Appendix.
We focus primarily on two families of data generating processes (DGPs) that embed different assumptions about the distribution of post-intervention potential outcome trends in the intervention and comparison groups: (1) group mean trends are different with overlap in the individual-level trends; and (2) group mean trends are different without overlap in the individual-level trends. Figure (ref) presents histograms of individual-level trends that illustrate the difference between the scenarios with and without overlap. In these DGPs, outcomes are a linear combination of the intervention effect, a deterministic linear trend with heterogeneous slopes across individuals, and an error term. Individual-level trends in potential outcomes absent treatment are constant over time, which means that Assumption (ref) is satisfied. Within each family of DGPs, we also explore the importance of reliability of the estimated pre-intervention outcome trends by simulating DGPs that vary in the degree of autocorrelation in the error term: DGPs with higher autocorrelation allow more reliable estimates of pre-intervention observed outcome trends.
Under each DGP, we compare the bias of three DD models that use different approaches for handling the differences in group mean pre-intervention outcome trends: an unadjusted approach, matching on the estimated linear trend, and entropy balancing on the estimated linear trend. For each of these approaches, we fit two DD models: one that includes time as a linear function (Equation (ref) with $P=1$) and one that includes time nonparametrically (Equation (ref)). For matching, we use 1:1 matching without replacement and a caliper of 0.2 standard deviations. For entropy balancing, we balance on the first moments of the estimated trends. We report estimator performance in terms of the percent bias reduction relative to the bias of the unweighted DD estimator: an unbiased estimator will have 100% bias reduction.
In results not shown here, we also specified a DGP such that the group counterfactual trends are the same. In this case, Assumption (ref) is satisfied without weighting and DD is unbiased without weighting. All DD estimators are approximately unbiased for all values of the residual autocorrelation. This analysis simply confirms that the weighted DD estimators considered here do not introduce bias in situations when unweighted DD is also unbiased.
First, we simulate DGPs that are ideal for balancing pre-intervention trends: there are mean differences in the counterfactual outcome trends, but there is overlap in the individual-level trends between the intervention groups (satisfying Assumption (ref)). The distribution of the slopes of the individual-level trends is illustrated in the left panel of Figure (ref), which confirms that all values of the intervention group are represented in the comparison group.
Figure (ref) reports the percent bias reduction comparing each estimator to the usual DD as a function of the residual autocorrelation, which varies from zero to 0.99. First, all four weighted DD estimators reduce bias. Second, given the balancing approach, estimators that control for linear trends (left panel) yield greater bias reduction than approaches that control for nonparametric time effects (right panel). This likely occurs because the true outcome trends are linear and are better approximated by the DD with time parameterized linearly. Third, for all estimators, the bias reduction increases as the residual autocorrelation increases. The pattern of improved bias reduction with increasing residual autocorrelation is likely to reflect the increasing reliability of the estimated outcome trends.\supercite{steiner2011importance} Fourth, entropy balancing on linear trends (left panel) removes most of the bias for all values of the residual autocorrelation. In contrast, a similar matching based estimator does not.
Here, we simulate DGPs that violate the assumption of overlap in the individual-level trends (Assumption (ref)). This is highlighted in the right panel of Figure (ref), where we have plotted the distribution of individuals' linear time trends by intervention group. Figure (ref) reports the percent bias reduction comparing each estimator to the usual DD. As in Figure (ref), bias reduction is plotted against the residual autocorrelation.
The general pattern of bias reduction is similar to that from the previous results, with two key distinctions. First, the balancing approaches do not remove all bias as the residual autocorrelation approaches 1. Second, the amount of bias reduction compared to the unadjusted model is less than that of the previous section. These results are due to the fact that there is no overlap in the trends, so the balancing approaches are unable to match the true pre-intervention trends even as the residual autocorrelation increases.
An unexpected finding is that entropy balancing of the linear outcome trends combined with a DD model that parameterizes time linearly is approximately unbiased for all values of the autocorrelation---even without overlap. We believe this reflects the double robustness property of entropy balancing, which states (among other things) that entropy balancing augmented with a regression model for the outcome is consistent, as long as the outcome model is correctly specified.\supercite{qingyuan2016} Here, we have correctly formulated the outcome regression as a linear model, potentially satisfying the assumption needed for consistency. Technically, our DD model is misspecified: the group mean trend differs between intervention groups, but the model assumes a common trend. However, the double robustness property for such models may require only correct specification of the outcome model among the comparison group due to the estimand of interest being defined as the ATT. More work is needed to theoretically justify this result for DD models.
To illustrate the use of entropy balancing in DD evaluation of a health policy intervention, we used data from the first and second annual evaluations of the MA VBID model test.\supercite{eibner2018first} VBID is an approach to health insurance design that tailors patient cost-sharing and other dimensions of benefit design on the basis of patients' health status. VBID aims to guide patients toward more appropriate utilization decisions, e.g., by reducing co-payments associated with health services or pharmaceuticals with a high clinical value given the patient's chronic conditions.\supercite{chernew2007value} Earlier demonstrations of VBID have shown promise in the employer-sponsored insurance market,\supercite{chernew2008impact,frank2012effect,choudhry2011full,choudhry2010pitney,choudhry2014five,gibson2011value,maciejewski2014value,yeung2017impact} but VBID had not previously been tested in Medicare beneficiaries aged 65 and over. Between 2017 and 2019, the Center for Medicare & Medicaid Innovation (CMMI) within the Centers for Medicare & Medicaid Services (CMS) conducted the first test of VBID in Medicare Advantage (MA), an intervention referred to as the MA VBID Model Test.\supercite{eibner2018first} In 2017, a total of 45 MA plans participated in the MA VBID model test. See eibner2018first\,\supercite{eibner2018first} for details on the implementation, benefit designs, and impacts of MA VBID, as well as details on data sources and data construction for the sample presented here.
In this paper, we provide DD estimates of the effect of VBID on per-beneficiary per year utilization of office-based primary care and specialty care. Table (ref) provides baseline characteristics and outcomes from 2016 of VBID-eligible beneficiaries in participating MA plans and their matched comparison beneficiaries in non-participating MA plans. Overall, baseline characteristics of beneficiaries are similar between VBID-participating and comparison MA plans, with beneficiaries in VBID plans being slightly older (77.8 vs 77.1) and having higher risk scores (1.80 vs 1.58). We apply entropy balancing using the full set of baseline characteristics from Table (ref) and the first-differences of the pre-intervention observed outcomes. After application of entropy balancing weights for both sets of outcomes, the means of the baseline characteristics from comparison beneficiaries match those of VBID beneficiaries exactly. Note that the outcome levels do not match, as the entropy balancing only seeks to balance pre-intervention outcome trends.
Figure (ref)(a) shows overlap between VBID and comparison groups in the beneficiary-level trend in the number of primary care visits from 2014 to 2016. Only a handful of the nearly 40,000 VBID-eligible beneficiaries lie outside the range of comparison beneficaries, suggesting sufficient overlap. Figure (ref)(b) shows the average number of primary care visits for VBID-eligible beneficiaries in VBID-participating plans and for comparison beneficiaries before and after entropy balancing. Trends from 2014 to 2016 were similar prior to entropy weighting, but comparison beneficiaries' utilization is increasing at a slower rate than VBID beneficiaries. A test of a departure of parallel trends prior to 2017 rejects the null hypothesis of parallel trends before entropy balancing ($p<0.001$), but not after ($p\approx 1$). This result is visually verified in Figure (ref)(b). Figures (ref)(a) and (ref)(b) show the same information for specialty visits, and similar conclusions are formed. A test of a departure of parallel trends prior to 2017 for specialty visits rejects the null hypothesis of parallel trends before weighting ($p<0.001$), but not after ($p\approx 1$). This result is visually verified in Figure (ref)(b).
The estimated increase in the number of primary care visits per beneficiary that is associated with VBID is 0.12 (0.08--0.17) before entropy balancing and 0.16 (0.11--0.21) after entropy balancing. A larger difference is observed for specialty care visits, with the increase in the number of visits associated with VBID estimated as 0.05 (-0.04--0.14) prior to entropy balancing and 0.14 ( 0.04--0.25) after entropy balancing---nearly three times larger an effect. In this example, DD after entropy balancing could lead to a substantially different conclusion about policy impacts than the conclusion one might draw from the unweighted DD estimate.
Our theoretical developments and simulation studies provide practical guidance on when it is beneficial to balance pre-intervention trends in the outcomes between intervention groups. Although we investigated only a few DGPs, we expect similar results to hold across many other settings. The simulation results presented in this paper were limited to cases where the counterfactual trends were linear, which guarantees that Assumption (ref) is satisfied. In the Appendix, we relaxed this assumption by simulating data with nonlinearities, and find that using nonparametric time simultaneously in entropy balancing and the DD model yields greater bias reductions when compared to balancing the estimated linear trends.
Based on the results of this study, we can offer some recommendations for health policy researchers. If the parallel trends assumption of the DD model is expected to hold for the group means, then there is no need to balance pre-intervention trends. However, entropy balancing of pre-intervention trends does not introduce bias. If there is uncertainty about whether the parallel trends assumption is expected to hold, then entropy balancing of pre-intervention outcome trends reduces bias relative to the unadjusted estimator. We also note that our theoretical and simulated results focus on outcome trends, and none of our findings justify balancing pre-intervention outcome levels: in this we agree with daw2018matching\,\supercite{daw2018matching}.
The amount of bias reduction that can be achieved using entropy balancing for DD, or any other similar approach, is limited by how well the individual-level outcome trends can be estimated. A highly reliable estimate of the individual-level trends will remove nearly all of the bias, while a completely unreliable estimate will remove no bias. The mechanism through which the estimated individual-level trends have higher reliability should not matter, and we expect any highly reliable estimated trend to remove considerable bias. Higher reliability is achieved as the autocorrelation increases, as the residual error decreases, as the strength of the time trends increases, as the individual-level heterogeneity of the time trends increases, or as the number of pre-intervention time points increases (if balancing a parametric outcome trend), among others. Adding additional individuals will not, in general, improve the reliability of the estimated outcome trends. Care should be taken to estimate the pre-intervention outcome trends in a manner that maximizes reliability without sacrificing robustness.
The assumptions presented in this paper are sufficient (not necessary) conditions for weighted DD to eliminate bias. It is possible that they can be relaxed or modified. In some applications, the parallel trends assumptions might be narrowed to specific time periods to account for anticipatory effects or a wash-out period. We also expect similar results to hold for other causal estimands identified using DD, such as the time-averaged ATT.
Entropy balancing for DD also offers a practical advantage for large-scale policy evaluations, such as the evaluation of MA VBID described in this paper, that analyze many outcome measures. Researchers may have difficulty identifying a comparison group that directly satisfies the parallel trends assumption uniformly across all outcomes. As in the MA VBID evaluation, entropy balancing on pre-intervention outcome trends can be used to refine a comparison group that can be used to analyze many different outcomes while reducing the bias of DD estimates when the parallel trends assumption is violated for specific outcomes. For instance, in the example presented in Section (ref), pre-intervention trends in the unweighted comparison group were much closer to balanced for primary care visits than for specialty care visits.
The use of DD to recover causal effects in health policy research and other settings rests on the untestable assumption that intervention and comparison groups have parallel post-intervention trends in potential outcomes absent the intervention, an assumption that is frequently evaluated using information about pre-intervention outcome trends. When pre-intervention outcome trends are not parallel between intervention and comparison groups, unweighted DD estimators are likely to be biased and should be avoided. The results of this paper shows that it is possible to reduce bias by combining DD with weighting on pre-intervention outcome trends, and in certain contexts obtain unbiased estimates of causal effects.
The key distinction between the analysis presented here and other recent work that reached more pessimistic conclusions about combining matching with DD \supercite{daw2018matching} is that we examined DGPs where heterogeneity in individual-level outcome trends can be found within the intervention and comparison groups. When there is common support, or overlap, between the intervention and comparison groups in the distribution of individual-level outcome trends, estimation of the microlevel DD after entropy balancing can greatly reduce the bias of estimates of the average treatment effect on the treated. We believe that such heterogeneity is likely to exist in many real-world settings of interest to health policy researchers. In these cases, our results provide a method for reducing bias in DD estimation even when pre-intervention group mean outcomes suggest a violation of the parallel trends assumption.
\printbibliography