Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,770 characters · 5 sections · 41 citation commands
(Visualizing) Plausible Treatment Effect Paths
JEL Codes: C12, C13
Keywords: dynamic treatment effects, post-selection inference, uniform inference
\thispagestyle{empty}
\setcounter{page}{1}
We are interested in the treatment effect path of a policy at discrete horizons $h = 1,...,H$. Examples include dynamic treatment effects in microeconomics, impulse response functions in macroeconomics, and event study paths in finance. We write $\beta=\{\beta_h\}_{h={1}}^{H}$ for the vector that collects this dynamic treatment effect path up to the fixed maximum horizon of interest $H$. We assume access to point estimates of the parameters $\beta_h$, denoted by $\hat{\beta}_h$, that correspond to the cumulative effect of the policy at horizon $h=1, \ldots H$. Throughout, we assume the vector that collects the estimated dynamic treatment effect path, $\hat\beta$, satisfies $\hat\beta \sim N(\beta,V_{\beta})$ and that we have access to the covariance matrix $V_{\beta}$. Leading examples to obtain such estimates include distributed lag models, local projections, and event studies.\footnote{We abstract away from approximation issues, but note that standard asymptotic approximations within these settings along with access to consistent asymptotic variance estimators motivate this setup.} We consider both point estimation and uncertainty quantification, though our focus will be on the latter. In particular, we introduce two approaches to visualize the uncertainty about the treatment path, which we call cumulative and restricted plausible bounds. Both bounds are often substantially tighter than traditional confidence intervals, and can provide useful insights even when traditional (uniform) confidence bands appear uninformative.
The standard approach in economics to quantify and visualize the uncertainty associated with parameter estimates is to construct confidence regions. Intuitively, a confidence region visualizes to the reader what values of the parameter, in this case $\beta$, are “plausible” based on the observed data. The idea being that values inside this region appear plausible, while values outside of the region do not. The two predominant confidence regions in practice are pointwise and sup-t confidence regions (e.g. callaway2021; jorda2023local; boxell2024). A third alternative is the Wald confidence region $CR^{Wald}$. This region simply collects all parameter values $b$ that are not rejected by a standard Wald test of the null hypothesis that $\beta = b$ at level $\alpha$. While a confidence region constructed from pointwise confidence intervals does not achieve correct coverage for the vector $\beta$, both sup-t and Wald confidence regions achieve valid coverage: $\mathbb{P}(\beta \in CR^{Wald})= \mathbb{P}(\beta \in CR^{sup-t}) = (1-\alpha)$.\footnote{We discuss these regions, and their construction, in more detail in Appendix (ref). For further discussion of uniform confidence bands, and, in particular, the merits of sup-t confidence bands, also see freyberger2018 and olea:pm.}
Sup-t and Wald confidence regions both come with some advantages and disadvantages. Since the Wald region is an ellipsoid, a disadvantage of the Wald confidence region is that it becomes infeasible to visualize in higher dimensions (i.e. when $H > 3$). The sup-t confidence region has the advantage of being easy to visualize. However, the volume of the sup-t confidence region quickly explodes relative to the volume of the Wald region. We illustrate this difference in confidence region volume in Figure (ref), which plots the volume of the Wald confidence region relative to the volume of the sup-t confidence region as a function of the dimension $H$ for two exemplary covariance matrices $V_{\beta}$. We see that the volume of the sup-t region tends to be orders of magnitude larger than the volume of the Wald region for the typical horizon that is depicted in event studies and impulse responses.\footnote{In what follows, we refer to the visualizations of $\hat{\beta}$ simply as treatment effect plots.} When $V_{\beta}$ is the identity matrix ($\rho=0$), the relative volume of the Wald region is less than 10% and around 0.1% of the volume of the sup-t region for $H=12$ and $H=24$ respectively. These numbers are generally even smaller if the entries in $\hat{\beta}$ have non-zero correlation: When $V_{\beta}$ is a symmetric Toeplitz matrix with entries $v_{ij}=0.95^{|i-j|}$, this ratio drops to 3.5% and 0.0001% for $H=12$ and $H=24$ respectively. One immediate consequence is that, for even moderate horizons $H$, the overwhelming majority of paths inside the sup-t bands would be rejected by a simple joint hypothesis test. This property seems unappealing to us and serves as a first indication that sup-t confidence bands may not always be appropriate for visualizing what dynamic treatment effect paths are plausible.
To illustrate this further, Figure (ref) depicts two exemplary treatment effect plots. The object of interest is the treatment path of a policy over the depicted horizon. We have access to jointly normal estimates $\{\hat{\beta}_h\}_{h=1}^{H}$, with observed point estimates $\hat{\beta}$ given by the black dots. Both panels further include pointwise 95 percent confidence intervals (inner confidence set as indicated by the dashes) and uniform 95 percent sup-t confidence bands (outer confidence set). While the pointwise confidence intervals only permit testing of pre-selected hypotheses for individual coefficients $\beta_h$, the sup-t bands contain the entire true path $\beta$ in 95 percent of realized samples. Figure (ref) depicts a hypothetical example with zero correlation between the estimated coefficients.\footnote{We give more detail on the underlying DGP in Section (ref).} Figure (ref) is based on the same estimates as Figure 4b in bosch2014. In this example, all off-diagonal entries in $V_{\beta}$ are positive, and the average correlation between adjacent coefficients is 0.95.
The sup-t region in Figure (ref) includes treatment paths that imply an overall positive effect (paths with $\sum_{h=1}^H \beta_h > 0$) and treatment paths with very different shapes. In fact, $\beta=0$ falls inside the sup-t bands, suggesting that the null of “no treatment effect" is plausible. However, a joint test of the null hypothesis that $\beta=0$ yields a p-value of $1.54 \times 10^{-9}$. In contrast, $\beta=0$ falls outside the sup-t bands in Figure (ref), suggesting that the null of “no treatment effect" is not plausible. However, a joint test of the null hypothesis that $\beta=0$ yields a p-value of $0.33$. These discrepancies between the easy-to-visualize sup-t region and the results of simple joint hypothesis tests again suggest to us that the sup-t confidence region may not always be providing an empirically effective visualization of what treatment effect paths are plausible.
In Figure (ref), we therefore introduce two alternative ways to visualize plausible treatment effect paths. Panels (ref)-(ref) are based on different simulated data generating processes, while panels (ref) and (ref) are based on published figures in Macroenomics (nakamura2018identification) and Applied Micro (bosch2014). Elements from the “standard” treatment effect plot are provided by the black dots which give point estimates of the treatment effect at each horizon and by the inner and outer confidence intervals corresponding respectively to the usual pointwise and sup-t confidence intervals.\footnote{Figure (ref) is based on the same estimates $\hat{\beta}$ as Figure (ref). Figure (ref) is based on the same estimates $\hat{\beta}$ as Figure (ref).} In addition, panels (ref)-(ref) include a solid blue line to represent the true treatment effect path and an additional dotted red line, which we explain below. Finally, each plot includes two new features: (i) the shaded red area, and (ii) the dashed and solid green lines. Importantly, these new features shift the goal posts relative to the Wald and sup-t bounds, and the inferential target of our bounds is not the true treatment path.
The shaded red area represents our proposed 95% cumulative plausible bounds. We construct these bounds so that the average treatment effect across the depicted horizons will be within these bounds for 95% of all realizations of the data. For example, in Figure (ref), these bounds suggest that the average effect of the policy over the 36 periods depicted is between (-0.248, -0.156), and thus that the overall effect of the policy over the 36 periods is strictly negative and inside the window (-8.93, -5.62). In contrast to the standard sup-t region, these bounds suggest that a treatment path with no overall effect of the policy is not plausible. These cumulative plausible bounds have an alternative interpretation in terms of the overall treatment effect path. Specifically, the cumulative plausible bounds are such that the treatment path $\beta$ (depicted as the solid blue line) will on average be within these bounds for 95% of all realizations of the data.
The dashed green lines represent our proposed 95% restricted plausible bounds, which are centered around restricted estimates provided by the solid green line. These restricted estimates and bounds are motivated by envisioning a researcher who is interested in understanding key features of the treatment effect path but is not concerned with necessarily covering the entire true path at every horizon. However, we also imagine the researcher as being ex ante unsure about what the important features are and wanting to use the data to help select a restricted model for summarizing the treatment effect path.
More concretely, we construct the restricted estimates and plausible bounds by using a statistical model selection procedure to select an approximating model from within a pre-specified universe of candidates. We consider a default set of models motivated by a preference for smooth dynamics that eventually die out induced by shrinking first and third differences of $\hat\beta$, though we note that the procedure could be applied with any finite, pre-specified universe of models. The restricted estimates are then simply the point estimates of the treatment path based on the selected model. We construct the restricted plausible bounds to provide uniform (95%) coverage accounting for data-dependent model selection by applying Berk et al.'s [berk2013valid] Post-Selection Inference (PoSI) to our setting.
Looking at the restricted estimates and restricted plausible bounds in each panel paints a starkly different picture compared to the sup-t intervals. In all cases, the restricted plausible bounds are relatively narrow and seemingly quite informative about the broad features of the treatment effect paths. Figure (ref) stands out and merits further discussion. In this instance, our model selection procedure selects a constant treatment effects model. Our restricted estimates then coincide with the MLE estimate of a constant treatment effects model.\footnote{That is, a model with $\hat\beta \sim N(\beta,V_{\beta})$, where $\beta$ is constant across $h$.} Remarkably, this estimate is $-0.0017$, which is outside of the convex hull of the individual estimates $\hat{\beta}_h$.\footnote{Intuitively, this behavior results from the strong positive correlation in the estimates combined with more precise estimates in early periods. We note that this strong positive correlation in the estimates cannot be inferred from the traditional plot.} Further, a Wald test of the null hypothesis that $\beta_h=-0.0017$ for $h=1, \ldots, H$ gives a p-value of 0.36. That is, a traditional joint hypothesis test suggests there is relatively little evidence against this hypothesis, which thus appears relatively “plausible,” contrary to what a visual inspection of the traditional treatment effect plot might suggest. This difference underscores that traditional treatment effect plots may be ineffective at visualizing the impact of the off-diagonal entries in $V_{\beta}$ (we illustrate this further in Appendix (ref)). In contrast, the entire covariance matrix $V_{\beta}$ is reflected in our restricted plausible bounds. Accounting for the covariance structure can lead to interestingly different results, improving the informativeness of these plots.
Finally, we reiterate that the inferential target of the restricted plausible bounds in the population is not the true treatment path. Rather, the restricted plausible bounds provide uniform coverage of a surrogate path given by the approximation that would be obtained by applying the selected model to the true effect path. We depict the selected surrogate for each of the simulated scenarios in Figure (ref) with a dotted red line. In Figure (ref) and Figure (ref), this surrogate is indistinguishable from the true treatment path. In Figure (ref) and Figure (ref), the surrogate differs from the true treatment path but visually captures what seem to be key features of the overall treatment path. Indeed, we suspect many empirical researchers, if given the true treatment path from Figure (ref), would actually be more interested in the smooth approximation provided by the surrogate in this case. That is, we view the fact that the restricted plausible bounds cover a data-dependent approximation to the population treatment effect path as a potentially appealing feature.
In summary, we propose augmenting standard event-study plots with two additional elements: a shaded red region (the cumulative plausible bounds) and dashed and solid green lines (the restricted plausible bounds and estimates, respectively). Together with the usual pointwise and sup-t intervals, these visualizations offer a more comprehensive view of plausible effect paths, each serving distinct inferential purposes. Sup-t bands provide a simple, assumption-free summary of plausible paths. Pointwise intervals target effects at specific horizons. Cumulative bounds inform average treatment effects, while restricted bounds and estimates capture approximations of the effect path obtained using data-driven smoothing.
Our paper connects to several strands in the literature. We obtain our restricted plausible bounds and estimates by considering a finite set of candidate models for $\beta$. This approach is thus closely related to work that considers parametric models and approximations to $\beta$. Such approaches have been studied going back at least to almon1965, who imposes a parametric model on distributed lag coefficients. More recently, barnichon2018functional propose approximating impulse responses with a set of basis functions, and barnichon2019 propose to shrink impulse response estimates towards polynomials. While related, we differ from these approaches by focusing on inference for a data-dependent surrogate effect path; see, for example, genovese2008adaptive for a general discussion of inference on surrogates.
We operationalize our restricted plausible bounds by using data-dependent selection from a universe of candidate models with different fixed degrees of shrinkage over first and third differences. This model universe is closely related to the structure employed in shiller1973 which takes a fully Bayesian approach to estimating a distributed lag model under a normal prior on the $d^{\textrm{th}}$ order difference of $\beta$. Our restricted estimates are thus akin to point estimates that could be obtained by taking an empirical Bayes approach within the framework of shiller1973. From the empirical Bayes perspective, one could then potentially adapt armstrong2022robust to the present context to obtain interval estimates.\footnote{Also see the SmIRF estimator of plagborg2016essays for a related approach that includes confidence sets with guaranteed frequentist coverage.} In contrast, our restricted plausible bounds provide frequentist coverage for the population value of the selected surrogate path.
To maintain coverage guarantees for the selected surrogate, accounting for data dependent model selection, we use a version of post-selection inference (PoSI) confidence intervals (berk2013valid). Given that our inferential target is the population value of the selected model, we note that one could adopt other approaches from the literature on selective inference; see, e.g., taylor2015statistical and kuchibhotla2022post for excellent reviews.
There are a variety of other approaches to quantifying and visualizing uncertainty about treatment effect paths available in the literature. For example, sims1999error argues that conventional pointwise bands common in the literature should be supplemented with measures of shape uncertainty, and proposes such measures. jorda2009simultaneous suggests a method to construct simultaneous confidence regions for impulse responses given propagation trajectories. freyberger2018inference propose a uniformly valid inference method for an unknown function or parameter vector satisfying certain shape restrictions. More generally, inference for the treatment effect path is tightly tied to more general nonparametric inference problems; see, e.g., chenadaptive2024 for an interesting recent example that explicitly allows for use of a data-dependent sieve dimension. The “shotgun plot” of inoue2016, which depicts a random sample of $B$ impulse responses contained in the joint Wald confidence set, provides an alternative approach to visualizing plausible treatment effect paths. We believe our proposal to provide simple additional visual elements to the usual treatment effect plot provides a useful complement to this existing literature.
We first present a simple visual feature, the cumulative plausible bounds, that can be added to a standard treatment effect plot. This visualization does not impose any functional form or smoothness assumptions on the underlying treatment path, but, in terms of the full treatment effect path, it also does not achieve uniform coverage. Rather, the cumulative plausible bounds use a weaker notion of “cumulative coverage": The true treatment path will on average be within the cumulative plausible bounds in ($1-\alpha$)% of all realizations of the data for a given significance level $\alpha$. That is, by providing valid inference for the average effect over the horizon $H$, our cumulative plausible bounds provide a simple visual element that conveys uncertainty about the average treatment effect.
These cumulative plausible bounds are simply visualizations of the dynamic treatment path corresponding to the largest and smallest sum of all treatment effects up to horizon $H$ not rejected by a standard hypothesis test. They can be interpreted as boundary paths that would be consistent with the upper and lower limits of a confidence interval for the overall effect of the policy over $H$ periods. Formally, let
where $\kappa^{(1-\alpha)}$ denotes the inverse of the chi-square cdf with one degree of freedom at chosen significance level $(1-\alpha)$. We further define $l^{1-\alpha}$ analogously, replacing the $\max$ in (ref) with $\min$. Since both $u^{1-\alpha}$ and $l^{1-\alpha}$, corresponding to the upper and lower limit of the overall effect are scalars, there are infinitely many treatment paths that correspond to these bounds on the overall treatment effect. To visualize the bounds, we use $(U,L)^{1-\alpha}=\{U^{1-\alpha}_h, L^{1-\alpha}_h\}_{h=1}^H$, where $U^{1-\alpha}_h = \frac{u^{1-\alpha}}{H}$ and $L^{1-\alpha}_h = \frac{l^{1-\alpha}}{H}$.\footnote{Instead of using bounds $(U,L)^{1-\alpha}$ that are constant across $h$, one could alternatively depict bounds that reflect the shape of the unrestricted estimates.} We choose this visualization as the interval $(U,L)^{1-\alpha}$ is a $(1-\alpha)$% Wald confidence interval for the average effect of the policy over the horizon $H$, $\frac{1}{H} \sum_{h=1}^H \beta_h$.
The following trivial proposition clarifies how coverage of these cumulative plausible bounds relates to coverage of the treatment path.
Proposition (ref) states that, for a given significance level $\alpha$, the true treatment path will on average be within our bounds for ($1-\alpha$)% of all realizations. This follows immediately from the fact that any path that is not, on average, inside the cumulative plausible bounds implies an overall treatment effect over $H$ periods that is rejected by the corresponding hypothesis test.
The second idea we pursue is to present confidence regions that cover approximations of the true effect path that have “reasonable shapes." We term these confidence regions restricted plausible bounds. Here, we define “reasonable shapes” by pre-specifying a universe of models. We then use data-dependent model selection to choose a good representation for $\hat{\beta}$ from among this set. Intuitively, this approach is related to directly imposing a functional form restriction as is often done in empirical work, for example by
One key feature of our approach is that we do not rely on a fixed functional form restriction or make use of some other implicit or ad hoc device to choose a restricted model. Rather, we select a model, and then take model selection explicitly into account when constructing confidence bounds. That is, we propose a model selection procedure that is explicit, transparent, and will allow us to maintain formal coverage guarantees instead of implicitly using the data to select a restricted model, which leads to invalid inference.
Before we formally define our proposal, we introduce some necessary notation. We first borrow from the nonparametric statistics literature to introduce the notion of a “surrogate" (cf. genovese2008adaptive). A surrogate path $\beta_M$ is close to, but potentially simpler than, $\beta$. We note that the surrogate path is a population object that approximates $\beta$, the true treatment path. For example, we may define a constant treatment effects surrogate of $\beta$ as $\beta_s=\operatorname*{arg\,min}_{b} \ (\beta - b)'(\beta - b) \ \text{ s.t. } \Delta b =0$. If the surrogate model $M$ is fixed a priori (and not itself a function of the data), inference for $\beta_M$ is straightforward, though we stress that any inferential statements in this case will be about $\beta_M$ and not $\beta$.\footnote{Targeting a simple surrogate function is akin to the standard approach in economics of estimating linear models even when the conditional expectation function is not believed to be linear. One can think of the linear model as a “surrogate model” capturing the best linear predictor. Inference will then be about the linear surrogate, and not the “truth.”} However, failing to take into account that the data is used to select the surrogate creates a problem for inference (e.g. leeb2005 or roth2022pretest). In our setting the surrogate is explicitly a function of the data (or more precisely, of the unrestricted estimates $\hat{\beta}$), and we may thus write $\beta_{M(\hat{\beta})}$ to denote a data dependent surrogate path. In a first step, we use the data to select the surrogate model. In a second step, we then create a uniformly valid confidence region for the selected surrogate path, taking into account that the choice of surrogate is also random (i.e. a function of the data).
Given that we are doing model selection from a specified universe of models, a key choice is the specific model universe we consider. We consider a model universe motivated by the following economic intuition:
In practice, we use shrinkage over first and third differences of $\hat\beta$ to implement 1. and 2.
Formally, we assume that the estimates of the treatment path $\hat{\beta}$ are jointly normal with $\hat{\beta} \sim N(\beta,V_{\beta})$, where $V_{\beta}= \sigma^2 V$, $\sigma^2=\frac{1}{H}\sum_{h=1}^H V_{\beta}(h,h)$, and $V$ is positive-definite. Taking $\hat{\beta}$ as input, we define the following object:
where
Solving (ref) provides a closed form solution\footnote{For intuition, note that the problem in (ref) is closely related to the following constrained optimization with tuning parameters $c_1$, $c_2$, and $K$, explicitly bounding the first and third difference:
However, this formulation does not have a closed form solution and is computationally more challenging, making it less appealing in practice.} for $\tilde\beta(\lambda_1,\lambda_2,K) := \tilde\beta(M)$ given by
For fixed $M = (\lambda_1, \lambda_2, K)$, it immediately follows that
where $V_{M}=P(M) V_{\beta}P(M)'$. Here, $P(M)\beta= \beta_M$ defines a particular surrogate path for $\beta$. Intuitively, $P(M)\beta$ corresponds to a “projection” of the true treatment path $\beta$ into a lower dimensional space. Given (ref), it would be straightforward to construct a confidence region for $\{\beta_{M,h}\}_{h=1}^H$, where $\beta_{M,h}$ denotes the $h^{\text{th}}$ entry in vector $\beta_{M}$, for a given, fixed value of the tuning parameters. However, knowing ex ante what values to use for $\lambda_1$, $\lambda_2$, and $K$ seems challenging. We thus use model selection to choose $\lambda_1$, $\lambda_2$, and $K$ --- or, equivalently, to choose the surrogate model $M$.
Specifically, we use the estimated $\hat\beta$ and an object akin to an information criterion to select the surrogate $M$. First, note that we can construct the “residuals" $\hat\beta - \tilde\beta(M) = \hat\beta - P(M)\hat\beta = (I-P(M))\hat \beta$. We use this residual formulation to define an analog of model degrees of freedom given by df$(M)$ = trace$(P(M))$. We then select a model that minimizes a BIC analog over $\mathcal{M}$ where $\mathcal{M}$ denotes the universe of values for $M=(\lambda_1,\lambda_2,K)$: $$ \hat{M} = \operatorname*{arg\,min}_{M \in \mathcal{M}} (\hat{\beta} - \tilde\beta(M))'V_{\beta}^{-1}(\hat{\beta} - \tilde\beta(M)) + \log(H) \text{df}(M).$$
We tie the researcher's hands by pre-specifying $\mathcal{M}$, the universe of models considered. In our implementation, $\mathcal{M}$ includes surrogate models corresponding to a constant, linear, quadratic and cubic treatment effect path (with one, two, three, and four degrees of freedom respectively), as well as an unrestricted model corresponding to the unrestricted estimates $\hat{\beta}$ (with $H$ degrees of freedom). $\mathcal{M}$ further includes surrogate models corresponding to surrogate paths of the form $P(M)\beta=\beta_M$ using a grid over $(\lambda_1,\lambda_2,K)$. We discuss our implementation in more detail in Appendix (ref) and visualize the model universe $\mathcal{M}$ for four exemplary treatment paths $\beta$ in Online Appendix Figure (ref), but note that $\mathcal{M}$ does not depend on $\hat\beta$ or $\sigma$.
Given the selected surrogate $\hat{M}$, we define the restricted estimates as $\tilde\beta(\hat{M})$. However, we cannot directly apply (ref) to obtain a valid confidence region for the population value of the surrogate path $\beta_{\hat{M}}=P(\hat{M})\beta$ because $\hat{M}$ was selected by looking at the data, $\hat\beta$. Thus, in a second step, we use Valid Post-Selection Inference (berk2013valid), which explicitly accounts for data-dependent (and thus random) model selection, to construct a uniformly valid confidence region for $\beta_{\hat{M}}$. These confidence intervals, our restricted plausible bounds, are rectangular regions of the form $CR^{POSI} = \{\ell_h(X), u_h(X)\}_{h=1}^H$ for $[\ell_h(X), u_h(X)] = [\tilde{\beta}(\hat{M})_h \pm C^{\alpha} V_{\hat{M}}^{1/2}(h,h)]$ where $\tilde{\beta}(\hat{M})_h$ denotes the restricted estimate of the effect at horizon $h$ and $V_{\hat{M}}^{1/2}(h,h)$ is the square root of the $h^{\text{th}}$ diagonal entry of $V_{\hat{M}}$. To ensure uniform validity we use the “PoSI constant” of berk2013valid as $C^\alpha$, defined as the minimal value that satisfies
where $t_{h \cdot M} = V_{M}^{-1/2}(h,h) \xi_h$, and $\xi_h$ is the $h^{th}$ element of multivariate normal vector $\xi$ with mean $\mathbf{0}_H$ and variance $V_M$. Importantly, $C^{\alpha}$ depends on $\mathcal{M}$, the universe of models considered, but not on the model selection procedure.
The following proposition is a direct application of berk2013valid.
Proposition (ref) guarantees that our restricted plausible bounds cover the selected surrogate to the truth in at least $(1-\alpha)$% of sample realizations.
In this section, we illustrate the properties of our restricted plausible estimates and bounds as well as our cumulative plausible bounds in simulation experiments with treatment paths generated to resemble treatment path dynamics that practitioners may encounter. We illustrate these treatment paths in Figure (ref). We consider a constant treatment effect path (cf. Figure (ref)); a treatment path that smoothly declines before flattening out after 17 periods (cf. Figure (ref)); a hump-shaped treatment path with dynamics that continue for the entire $H$ periods (cf. Figure (ref)); and a “wiggly” treatment path (cf. Figure (ref)). We describe the exact DGP for each of the four panels in more detail in Online Appendix Table (ref).\footnote{Figure (ref) provides one example realization from each of these DGPs.} The object of interest is the treatment path over a 36-month horizon. We have access to jointly normal estimates $\{\hat{\beta}_h\}_{h=1}^{36}$. For each of these four treatment paths, we then draw 1,000 realizations of $\hat{\beta} \sim N(\beta,V_{\beta})$.\footnote{In the figures that follow, $V_{\beta}$ is diagonal with its entries specified in Online Appendix (ref). We repeat our exercise with more general covariance matrices in Online Appendix (ref).}
We first compare the point estimation properties of the unrestricted estimates $\hat{\beta}$ with our restricted estimates $\tilde{\beta}(\hat{M})$ for each of these four scenarios. In particular, Figure (ref) depicts the ratio in mean-squared error, $MSE_{\tilde{\beta}(\hat{M})}/MSE_{\hat{\beta}}$, as a function of $\sigma^2$, which scales the covariance matrix of the estimates, $V_\beta$ (see Online Appendix (ref) for more detail).
The largest value of $\sigma^2$ in Figure (ref) (corresponding to the left most point) thus represents relatively noisy estimates.\footnote{Figure (ref) was created from a single realization of the the left most point in Figure (ref). To give the reader a sense of the scale of the x-axes, Online Appendix Figure (ref) also illustrates a single realization of the right most point in Figure (ref).} We conclude that in most cases our restricted estimator has excellent point estimation properties when the target is the true treatment path $\beta$. In the three panels that have a smooth treatment path (Figures (ref)-(ref)), the MSE of our restricted estimate is a full order of magnitude lower compared to the unrestricted estimate. In Figure (ref), our restricted estimate has a lower MSE when estimates are very noisy, a higher MSE for intermediate sizes of $\sigma^2$, and a similar MSE when $\sigma^2$ is small. However, we suspect that a lower-dimensional summary of $\beta$, as provided by the surrogate path, may in fact be the policy relevant object in cases where the true treatment path exhibits complicated dynamics as in this panel (cf. Figure (ref)).
We next turn our attention to inference and first look at coverage. We again consider the four previous DGPs and vary the amount of noise in the estimation of $\beta$ by varying $\sigma^2$. The result is depicted in Figure (ref), where we set $\alpha=0.05$.
The pointwise, sup-t, and restricted coverage numbers, as indicated by the three solid lines, represent the empirical analogue to the usual notion of uniform coverage for $\beta$: It reflects the fraction of simulations in which the true treatment path $\beta$ falls entirely inside the pointwise confidence intervals (black), sup-t intervals (red), and restricted plausible bounds (yellow), respectively.
Across all four DGPs, the pointwise confidence intervals only cover the true treatment path in 15-20% of simulations. On the other hand, the sup-t intervals achieve nominal coverage across DGPs. The restricted plausible bounds are not constructed to provide uniform coverage for the true treatment path. It is thus not surprising that these bounds do not cover the entire treatment path in 95% of simulations across all DGPs. However, we do see that their coverage appears to converge towards 95% as the amount of noise in the initial estimates decreases (cf. Remark (ref)). Further, the restricted plausible bounds perform substantially better than the pointwise confidence intervals in covering the true treatment path when the true path is smooth. In the scenario of a wiggly treatment path, the restricted bounds exhibit poor coverage properties when the amount of noise in the initial estimates is large.
Proposition (ref) guarantees that the plausible bounds cover the selected surrogate to the truth in at least 95% of realizations. In Figure (ref), the restricted plausible bounds' coverage of the surrogate is illustrated by the dashed purple line. We see that this coverage is above 95% for all levels of $\sigma^2$ across all DGPs. We reemphasize that, in the scenario of a wiggly treatment path, a smooth approximation to the true path, for which we obtain valid coverage, may be a policy relevant object.
Finally, we note that, in line with Proposition (ref), the cumulative bounds (as indicated by the dashed green lines) indeed cover the cumulative effect of the policy in 95% of simulations. That is, the true treatment path is on average within these bounds in 95% of simulations.
In order to cover the true path in a higher fraction of simulations, the sup-t bands are generally wider than the pointwise bands (cf. Figure (ref)). We illustrate this across our simulations in Figure (ref), which depicts the average width of the sup-t bands and the restricted bounds, relative to the pointwise intervals.
While the sup-t bands are wider than the pointwise bands across the parameter space, the restricted bounds are in fact narrower throughout much of the parameter space, despite their improved coverage properties. For example, in the “Constant Treatment Effect" scenario, the restricted bounds are less than half as wide as the pointwise bands and less than a quarter as wide as the sup-t bands for all noise levels $\sigma^2$. Yet, they cover close to 95% of all realizations, while the pointwise confidence intervals only cover the true treatment path in 15-20% of simulations.\footnote{As we show in Online Appendix Figure (ref), we select surrogate paths $\beta_M$ with significantly reduced degrees of freedom, relative to the unrestricted estimates, in the first three DPGs. This effective dimension reduction explains both the improvement in point estimation properties of $\tilde{\beta}(\hat{M})$ relative to the unrestricted estimates $\hat\beta$ documented in Figure (ref), and the decrease in width of the confidence intervals documented in Figure (ref).} The restricted bounds are only wider than the sup-t bands when (i) the treatment path is wiggly and (ii) there is very little noise. In this region of the parameter space, our model selection mechanism infers that the wiggles in the treatment path are large enough relative to the noise level that heavy smoothing will incur a large cost in terms of losing fit. With relatively little smoothing, our restricted plausible bounds are then slightly more conservative than the standard sup-t bands, which follows immediately from the results in berk2013valid.
Finally, we shed some more light on the model selection of our procedure in Figure (ref).
Each panel depicts the 1,000 selected surrogates across our simulations in light brown. Since the chosen surrogates will depend on the precision of the initial estimates $\hat{\beta}$, we fix $\sigma^2$ at $0.014$, corresponding to the amount of noise in the examples in Figure (ref). When the true treatment path is smooth (Panels (ref)-(ref)), the chosen surrogate tends to closely approximate the truth. In contrast, the chosen surrogate is often simpler than $\beta$ while retaining its main features when the true treatment path is not smooth (Panel (ref)).\footnote{In Online Appendix Figure (ref), we further illustrate how the chosen surrogates depend on $\sigma^2$ for our Wiggly DGP.}
We are interested in (joint inference on) the dynamic treatment effect of a policy. We propose two visualization devices, which we term restricted and cumulative plausible bounds, to include in standard visualizations of estimated treatment effect paths. Because both bounds explicitly target objects other than uniformly covering the entire treatment path, they can be substantially tighter than standard pointwise and uniform confidence bands. Our bounds may thus provide useful insights about features of treatment effect paths when traditional confidence bands appear uninformative. Our bounds can also lead to markedly different conclusions relative to traditional visualizations when there is significant correlation in the estimates. As an auxiliary benefit, producing our restricted plausible bounds also provides additional restricted estimates that capture simplified representations of the dynamic effect path. Unsurprisingly, these restricted estimates have improved point estimation properties relative to the unrestricted estimates in settings where the treatment effect path is smooth.
It may be interesting to explore the notion of explicitly covering data-dependent surrogates in other economically relevant settings. While alternative model universes or model selection rules may be needed outside of the present setting, our proposal can easily be adapted to do so. For example, it would be interesting to explore related ideas in more structural settings.
Online Appendix