Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,117 characters · 13 sections · 71 citation commands
Doubly Robust Uniform Confidence Bands for Group-Time Conditional Average Treatment Effects in Difference-in-Differences
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} difference-in-differences, dynamic treatment effects, event study, treatment effect heterogeneity, panel data
\spacingset{1.8}
Difference-in-differences (DiD) is a powerful quasi-experimental approach to estimate meaningful treatment parameters. The recent DiD literature predominantly contributes to the development of identification and estimation methods in the staggered adoption case, where each unit continues to receive a binary treatment after the initial treatment receipt. callaway2023difference, de2023survey, roth2023s, and sun2022linear review recent contributions.
\phantomsection\Copy{AE-2-1}{ In empirical research using the DiD method, it is essential to understand the heterogeneity in treatment effects with respect to covariates, as well as “groups” and periods. As a concrete empirical example, suppose that we are interested in assessing whether and how minimum wage increases reduce poverty using county-level panel data. In this scenario, it is important to understand how the instantaneous and dynamic effects of minimum wage increases on the unemployment rate depend on the “pre-treatment” poverty rate. For example, if minimum wage increases result in significant lasting job losses but no substantial wage gains in high-poverty counties, then minimum wage increases would not effectively reduce poverty. In this case, policymakers should explore alternative policies for reducing poverty instead of relying heavily on the minimum wage policy. }
\phantomsection\Copy{AE-3-1}{ In this paper, we develop identification, estimation, and uniform inference methods to examine the treatment effect heterogeneity with respect to covariate values and other key variables (i.e., groups, periods, and treatment exposure time) in the staggered DiD setting. } We build on the setup of callaway2021difference and consider two types of target parameters: (i) the group-time conditional average treatment (CATT) function given a continuous pre-treatment covariate of interest and (ii) a variety of summary parameters that aggregate CATTs with certain estimable weights. We begin by showing that, under essentially the same identification conditions as in callaway2021difference, CATT is identifiable from a conditional version of the doubly robust (DR) estimand in callaway2021difference. Then, we propose three-step procedures for estimating CATT and the summary parameters: the first stage is the same as the parametric estimation procedures for the outcome regression (OR) function and the generalized propensity score (GPS) in callaway2021difference; the second and third stages comprise nonparametric local polynomial regressions (LPR) for estimating certain nuisance parameters and the conditional DR estimand. Lastly, to construct uniform confidence bands for the target parameters, we develop two uniform inference methods based on an analytical distributional approximation result and weighted/multiplier bootstrapping.
We investigate two statistical properties of our methods under the asymptotic framework where the number of cross-sectional units is large and the length of the time series is small and fixed. First, we derive asymptotic linear representations and asymptotic mean squared errors (MSEs) of our estimators, which are used for constructing standard errors and for choosing appropriate bandwidths. This part of the asymptotic investigations builds on the theory of the LPR estimation (fan1996local). Second, we prove uniformly valid distributional approximation results for studentized statistics and their bootstrap counterparts, which play an essential role in constructing asymptotically valid critical values for the uniform confidence bands. This result extends lee2017doubly and fan2022estimation, who study the uniform confidence bands for the conditional average treatment effect (CATE) function in the so-called unconfoundedness setup, to the staggered DiD setting. More precisely, similar to these prior studies, \phantomsection\Copy{AE-17-2}{ we use approximation theorems for suprema of empirical processes and asymptotic theory for potentially non-Donsker empirical processes by chernozhukov2014anti,chernozhukov2014gaussian to prove a uniformly valid analytical distributional approximation result and the uniform validity of weighted/multiplier bootstrap inference. }
Our key assumptions are, in line with callaway2021difference, the staggered treatment adoption and the conditional parallel trends assumption for panel data, in contrast to the unconfoundedness assumption for cross-sectional data. \phantomsection\Copy{AE-3-2}{ Unlike the unconfoundedness approach proposed by lee2017doubly and fan2022estimation, our methods allow researchers to learn about the heterogeneity of treatment effects with respect to key variables specific to the staggered DiD, such as groups, calendar time, and elapsed treatment time, as well as covariate values. } This attractiveness of our proposal is achieved with identification, estimation, and uniform inference methods tailored to the staggered DiD, whose statistical properties are not trivial from the existing results. \phantomsection\Copy{AE-1a-1}{ In particular, since our CATT is a causal parameter that captures the conditional average treatment effect on the treated, we need to estimate nonparametric nuisance parameters in the second-stage estimation, and the construction of our uniform confidence bands requires careful consideration of its impact on conditional DR estimates. } \phantomsection\Copy{AE-3-3}{ In addition, we highlight the importance of selecting appropriate critical values and bandwidths to ensure the uniform validity not only over covariate values but also over other key variables. }
\paragraph{Related Literature.}
This paper focuses on CATT and the conditional aggregated parameter in the staggered DiD setup as the target parameters, building directly on the focus on ATT and the unconditional aggregated parameter by callaway2021difference. In doing so, we develop uniformly valid inference for treatment effect heterogeneity with respect to covariate values and other key variables while taking advantage of callaway2021difference, that is, multiple treatment timing, treatment effect heterogeneity across units and time, the useful aggregation, and the attractive DR property. While all of these are empirically desirable, the DR property should be particularly important when performing uniform inference for treatment effect heterogeneity with respect to covariate values, as highlighted by lee2017doubly and fan2022estimation in the unconfoundedness setup. This is the main reason why we build directly on callaway2021difference rather than the other useful DiD methods (e.g., sun2021estimating,wooldridge2021two,borusyak2024revisiting).
In terms of developing uniform inference for treatment effect heterogeneity with respect to covariate values, we build on directly lee2017doubly and fan2022estimation in the unconfoundedness setup. Compared to their cross-sectional data analyses, our panel data analysis allows for understanding both the static and dynamic nature of treatment effect heterogeneity, which should be important from an empirical point of view. From a theoretical perspective, this advantage of our proposal stems from our DiD approach with (i) (possibly long) time differences of outcomes rather than their levels and (ii) novel DR estimators constructed with parametric and nonparametric nuisance function estimators and weighting schemes different from theirs. In particular, estimating the nonparametric nuisance functions in our second-stage estimation has non-negligible (first-order) effects on the asymptotic properties of our DR estimator, as shown in Theorems (ref) and (ref), which are new insights in the literature. \phantomsection\Copy{AE-1a-2}{ The issue of estimating nonparametric nuisance functions arises in our case because our CATT is a type of the conditional average treatment effect on the treated, as opposed to the focus on CATE in lee2017doubly and fan2022estimation. } After dealing with these considerations by building on the theory of the LPR estimation, our uniformly valid approximation results in Theorems (ref) and (ref) are obtained as applications of empirical process techniques in the same manner as in these two previous studies.
This paper clearly builds on previous work in the DiD literature that developed methods to understand treatment effect heterogeneity arising from covariates. For example, abadie2005semiparametric proposed pointwise inference for the conditional average treatment effect on the treated given a covariate based on inverse probability weighting (IPW) and series approximations in the canonical two-periods and two-groups DiD setting. His proposal is even applicable to the staggered adoption case by focusing on a subset of the original dataset consisting only of a treated group with a specific treatment timing and a comparison group. Compared to his proposal, our methods have novelties in terms of the empirically desirable DR property, the kernel smoothing technique that facilitates tuning parameter selection, and the uniform validity over covariate values and other key variables proven by empirical process theory.
\paragraph{Paper Organization.} The rest of the paper is organized as follows. Section (ref) introduces the setup and provides a non-technical roadmap for implementing our methods. Section (ref) illustrates our methods in the context of the minimum wage. Sections (ref) and (ref) discuss the identification, estimation, and uniform inference methods for CATT and the summary parameters, respectively. The supplementary results are presented in the online appendix. The accompanying R package { \href{https://tkhdyanagi.github.io/didhetero/}{didhetero}} is available from the authors' websites.
Whenever possible, we use the same notation as in callaway2021difference. For each unit $i \in \{ 1, \dots, n \}$ and time period $t \in \{ 1, \dots, \mathcal{T} \}$, we observe a binary treatment $D_{i,t} \in \{ 0, 1 \}$, an outcome variable $Y_{i,t} \in \mathcal{Y} \subseteq \mathbb{R}$, and a vector of the pre-treatment covariates $X_i \in \mathcal{X} \subseteq \mathbb{R}^k$. For notational simplicity, we often suppress the subscript $i$.
We consider the staggered adoption design, which includes the canonical two-periods and two-groups setting as a special case, under the random sampling scheme for balanced panel data.
Denote the time period when the unit becomes treated for the first time as $G \coloneqq \min\{ t : D_t = 1 \}$. We set $G = \infty$ if the unit has never been treated. We often refer to $G$ as the “group” to which the unit belongs. In particular, we call the set of units with $G = g$ for $g \in \{ 2, \dots, \mathcal{T} \}$ as the “not-yet-treated” group in pre-treatment periods $t < g$ and that with $G = \infty$ as the “never-treated” group. Assuming that $\bar g \coloneqq \max_{1 \le i \le n} G_i$ is known a priori, we write the set of realized treatment timings before $\bar g$ as $\mathcal{G} \coloneqq \mathrm{supp}(G) \setminus \{ \bar g \}$. With an abuse of notation, we let $\bar g - 1 = \mathcal{T}$ if $\bar g = \infty$.
Under Assumption (ref), the potential outcome given $G$ is well-defined. Specifically, we write $Y_t(g)$ as the potential outcome in period $t$ given that the unit becomes treated at period $g \in \{ 2, \dots, \mathcal{T} \}$. Meanwhile, we denote $Y_t(0)$ as the potential outcome in period $t$ when the unit belongs to the never-treated group (i.e., when $G = \infty$). By construction, $Y_t = Y_t(0) + \sum_{g = 2}^{\mathcal{T}} [ Y_t(g) - Y_t(0) ] \cdot G_g$, where $G_g \coloneqq \bm{1}\{ G = g \}$. Note that $Y_t(g) - Y_t(0)$ is the effect of receiving the treatment for the first time in period $g$ on the outcome in period $t$.
We aim to examine the extent to which the average treatment effect varies with groups, periods, and a single continuous covariate. To be specific, suppose that $X$ can be decomposed into $X = (Z, X_{\operatorname*{\mathrm{sub}}}^\top)^\top$ with a scalar continuous covariate $Z$ and the other elements $X_{\operatorname*{\mathrm{sub}}}$. The presence of $X_{\operatorname*{\mathrm{sub}}}$ should be important in typical DiD applications where the parallel trends assumption is more likely to hold only after conditioning on a number of covariates. For some pre-specified real numbers $a$ and $b$ such that $a < b$, let $\mathcal{I} = [a, b]$ denote a proper closed subset of the support of $Z$. As the first target parameter, we consider the group-time conditional average treatment effect (CATT) given $Z = z$ for $z \in \mathcal{I}$:
Estimating ${\operatorname*{\mathrm{CATT}}}_{g,t}(z)$ over $(g,t,z)$ is helpful in understanding the treatment effect heterogeneity with respect to group $g$, calendar time $t$, and covariate value $z$.
In Section (ref), we develop identification, estimation, and uniform inference methods for CATT. We begin by introducing a conditional DR estimand $\operatorname*{\mathrm{DR}}_{g,t}(z)$, which is a conditional counterpart of the DR estimand in callaway2021difference, by using the not-yet-treated group as the comparison group. We then show that $\operatorname*{\mathrm{CATT}}_{g,t}(z)$ is identified by $\operatorname*{\mathrm{DR}}_{g,t}(z)$ for each $(g, t, z) \in \mathcal{A}$, where
Given the identification result, we propose to construct a $(1 - \alpha)$ uniform confidence band for $\operatorname*{\mathrm{CATT}}_{g,t}(z)$ over $(g,t,z) \in \mathcal{A}$ by a family of intervals, denoted as $\mathcal{C} \coloneqq \{ \mathcal{C}_{g,t}(z): (g,t,z) \in \mathcal{A} \}$ with
where $\widehat{\operatorname*{\mathrm{DR}}}_{g,t}(z)$ is a three-step estimator computed with certain parametric estimation procedures and nonparametric LPR estimation, $\widehat{\operatorname*{\mathrm{SE}}}_{g,t}(z)$ is a pointwise standard error, and $c(1 - \alpha)$ is a uniform critical value obtained from an analytical method or weighted bootstrapping. \phantomsection\Copy{AE-3-4}{ Importantly, to ensure that the uniform confidence band $\mathcal{C}$ is uniformly valid over $(g, t, z) \in \mathcal{A}$, the critical value $c(1 - \alpha)$ must not depend on $(g, t, z)$ and is larger than the standard Wald-type pointwise critical value (i.e., the $(1 - \alpha/2)$ quantile of the standard normal distribution). As will be discussed later, the bandwidth used for the LPR estimation is crucial for constructing $c(1 - \alpha)$, and we recommend using a bandwidth that does not depend on $(g, t, z)$ for our uniform inference. See Remark (ref). }
We can also consider a variety of useful summary parameters by aggregating $\operatorname*{\mathrm{CATT}}_{g, t}(z)$'s. Specifically, building on the aggregation scheme of callaway2021difference, we set the second target parameter to the aggregated parameter of the following form:
where $w_{g,t}(z)$ is a known or estimable weighting function that determines the causal interpretation of $\theta(z)$. For example, letting $e = t - g \ge 0$ denote elapsed treatment time, we can consider the “event-study-type” conditional average treatment effect:
This is the conditional counterpart of the event-study-type summary parameter in callaway2021difference and useful for understanding the treatment effect heterogeneity with respect to treatment exposure time $e$ and covariate value $z$. Another useful example is the simple weighted conditional average treatment effect, which aggregates $\operatorname*{\mathrm{CATT}}_{g,t}(z)$'s into an overall effect as follows:
where $\kappa(z) \coloneqq \sum_{g \in \mathcal{G}} \sum_{t=2}^{\mathcal{T}} \bm{1}\{ g \le t < \bar g \} \cdot \Pr(G = g \mid G < \bar g, Z = z)$. We can also consider other useful summary parameters by appropriately choosing different weights. See Appendix (ref).
In Section (ref) and Appendix (ref), we study how to perform the uniform inference for these summary parameters. The proposed uniform confidence band for the aggregated parameter $\theta(z)$ has the same form as the uniform confidence band for CATT, namely $\mathcal{C}_{\theta} \coloneqq \{ \mathcal{C}_{\theta}(z) \}$ with
where $\widehat{\theta}(z)$ is an estimator obtained as an empirical analogue of (ref), $\widehat{\operatorname*{\mathrm{SE}}}_{\theta}(z)$ is a pointwise standard error, and $c_{\theta}(1 - \alpha)$ is a uniform critical value via an analytical method or multiplier bootstrapping. \phantomsection\Copy{AE-3-5}{ Similar to the case of CATT, the uniform critical value $c_{\theta}(1 - \alpha)$ and the bandwidth should not depend on $z$ and the variable specific to the summary parameter of interest (e.g., treatment exposure time $e$). }
\phantomsection\Copy{AE-16-1}{ To justify the uniform confidence bands (ref) and (ref), we need to make bias arising from kernel smoothing asymptotically negligible. To this end, we propose an undersmoothing approach based on the insight of the simple robust bias-corrected (RBC) inference. Specifically, we consider estimating the target parameters by local quadratic regressions (LQR) based on integrated mean squared error (IMSE) optimal bandwidths for local linear regressions (LLR). See Section (ref). }
\phantomsection\Copy{AE-7-1}{ In the same spirit of the focus on “pre-trends” in previous studies, it would be beneficial to assess the credibility of the identifying assumptions using our uniform inference. To this end, we focus on the following testable implications in the pre-treatment periods: $\operatorname*{\mathrm{CATT}}_{g,t}(z) = \theta_{\operatorname*{\mathrm{es}}}(e, z) = 0$ for all $g \in \mathcal{G}$, $t \ge 2$ such that $t \le g - 2$, $e \le -2$, and $z \in \mathcal{I}$. Note that we exclude $t = g - 1$ and $e = -1$ as base periods. If we find estimation and uniform inference results that are inconsistent with these testable implications, it suggests violations of the identifying assumptions. We discuss this type of simple diagnosis based on pre-trends in Appendix (ref). }
Throughout the paper, we focus on the treatment effect heterogeneity with respect to a single continuous covariate $Z$, rather than the full covariate $X$. This is because focusing on a single continuous covariate of interest allows us to easily visualize and interpret the heterogeneity in instantaneous and dynamic treatment effects with respect to its values, as illustrated in Figure (ref) in the next section.
In this section, we use our proposal to assess the heterogeneity in the effects of the minimum wage change on youth employment. In doing so, we illustrate the empirical relevance of our methods using a real dataset before proceeding to the technical discussions in the following sections.
We use the same dataset as in callaway2021difference, which includes county level minimum wages, county level teen employment, and other county characteristics for 2,284 U.S. counties in 2001-2007. The outcome variable $Y_{i,t}$ is the logarithm of teen employment in county $i$ at year $t$. We define the group $G_i$ by considering 100, 223, and 584 counties that increased their minimum wages in 2004, 2006, and 2007, respectively, as the treated units. This implies that the remaining 2,184, 1,961, and 1,377 counties in 2004, 2006, and 2007, respectively, are the not-yet-treated units in each year. In line with the specification in the empirical analysis of callaway2021difference, the pre-treatment covariates $X_i$ consist of county characteristics before 2000, including the poverty rate (i.e., the share of the population below the poverty line), the share of the white population, the share of the population of high school graduates, the regional dummy, the median income, the total population, and the squares of the median income and the total population. To save space, we relegate more information about the dataset, summary statistics, and pre-trends to Appendix (ref).
\phantomsection\Copy{AE-2-2}{ Among the covariates, we focus on examining the treatment effect heterogeneity with respect to the poverty rate. Reducing poverty should be one of the main purposes of the minimum wage policy, but it is unclear a priori whether and how minimum wage increases reduce poverty. This is because the extent to which minimum wage increases reduce poverty depends on the structural relationship between wage gains and job losses at the bottom of the income distribution, as well as other factors that determine income, as discussed in dube2019minimum. From this perspective, understanding how the impact of minimum wage increases on teen employment depends on the poverty rate should be useful for assessing whether the minimum wage policy alleviates poverty. For example, if we find that minimum wage increases significantly decrease teen employment in low-poverty counties, but have no significant effect in high-poverty counties, policymakers should emphasize the importance of the minimum wage policy for poverty reduction at least in high-poverty counties. }
\phantomsection\Copy{AE-6-1}{ In Figure (ref), panel (a) shows the estimation and uniform inference results for ${\operatorname*{\mathrm{CATT}}}_{g,t}(z)$ for $(g, t) \in \{2004, 2006, 2007\}^2$. We restrict our focus to this set of $(g, t)$ for presentation purposes. Panels (b) and (c) depict the event-study-type conditional average treatment effect $\theta_{\operatorname*{\mathrm{es}}}(e, z)$ for $e \in \{ 0, 1, 2, 3 \}$ and the simple weighted conditional average treatment effect $\theta_{\operatorname*{\mathrm{W}}}^{\operatorname*{\mathrm{O}}}(z)$, respectively. Panels (b) and (c) use data from all available groups and post-treatment periods, not just data from $(g, t) \in \{2004, 2006, 2007\}^2$. } In each panel, the horizontal axis corresponds to the interval $\mathcal{I}$ set as the interquartile range of the poverty rate, the solid line indicates the LQR estimates based on the IMSE-optimal bandwidth for the LLR estimation, and the gray area corresponds to the 95% uniform confidence band via weighted/multiplier bootstrapping using mammen1993bootstrap's (mammen1993bootstrap) weights. \phantomsection\Copy{AE-3-6}{ Because the bandwidth used for the LQR estimation should not depend on the variables of interest (e.g., $(g, t, z)$ for $\operatorname*{\mathrm{CATT}}_{g,t}(z)$) for our uniform inference, we take the minimum of the integrated (over $z \in \mathcal{I}$) MSE-optimal bandwidths across the variables (e.g., groups $g$ and post-treatment periods $t$). } \phantomsection\Copy{AE-9-1}{ We also found that the LLR-based inference methods (both analytical and bootstrap) and the LQR-based analytical method lead to almost the same empirical results as those presented here, but we suppress them to save space. }
The main empirical findings can be summarized as follows. First, the estimated CATT functions are nearly flat around zero for the 2004, 2006, and 2007 groups in 2004, 2006, and 2007, respectively, and the corresponding uniform confidence bands are not as wide. This means that, with satisfactory precision in terms of uniform inference, minimum wage increases have on average almost no instantaneous effect on teen employment. Second, we find negative but small CATT estimates with a small amount of treatment effect heterogeneity for the 2004 and 2006 groups in 2006-2007 and 2007, respectively, and the corresponding uniform confidence bands are wider than those for the instantaneous effects. This result may suggest that there are small but not substantial dynamic effects of minimum wage increases on teen employment, but this may be due to the lack of precision of uniform inference for dynamic effects. Third, the uniform inference results for the summary parameters also imply that there are no substantial effects of minimum wage increases on teen employment with modest treatment effect heterogeneity. \phantomsection\Copy{AE-2-3}{ In particular, the result for the simple weighted conditional average treatment effect indicates that there is almost no effect, especially at high poverty rates. Overall, our empirical results suggest that minimum wage increases do not substantially decrease teen employment and thus may be effective in reducing poverty, particularly in high-poverty counties. }
In this section, we develop identification, estimation, and uniform inference methods for CATT defined in (ref). For this purpose, there are two options for the comparison group: the not-yet-treated group and the never-treated group. For presentation purposes, the main body of the paper presents only the analysis using the not-yet-treated group. The analysis using the never-treated group is relegated to Appendix (ref). Moreover, we can consider three types of estimands: OR, IPW, and DR estimands. \phantomsection\Copy{AE-10}{ Throughout the paper, we focus on the DR estimand as it is more robust against model misspecifications.}
The following quantities play important roles in the analysis using the not-yet-treated group and the DR estimand. For each $g$ and $t$, we define the generalized propensity score (GPS) and the OR function respectively by
Let $p_{g,t}(X; \pi_{g,t})$ and $m_{g,t}(X; \beta_{g,t})$ be parametric specifications for these quantities, where each function is known up to the corresponding finite-dimensional parameter. Denote the corresponding parameter spaces as $\Pi_{g,t}$ and $\mathscr{B}_{g,t}$.
We impose the following identification conditions, which are essentially the same as Assumptions 3, 4, 6, and 7(iii) in callaway2021difference.
The conditional DR estimand based on the not-yet-treated group is defined by
where $W \coloneqq (Y_1, \dots, Y_\mathcal{T}, X^\top, D_1, \dots, D_\mathcal{T})^\top$ and
As a building block of our uniform inference methods, the next lemma shows that our DR estimand identifies CATT if at least one of the GPS and OR function is specified correctly. This result follows from almost the same arguments as in Theorem 1 of sant2020doubly and Theorem 1 of callaway2021difference. In fact, the only difference is that our estimand is conditioned on a single covariate $Z$, while their estimands integrate over all covariates.
We develop estimation and uniform inference methods for CATT based on the identification result in Lemma (ref). With an abuse of notation, we write $m_{g,t} \coloneqq m_{g,t}(X; \beta_{g,t}^{*})$ and $R_{g,t} \coloneqq R_{g,t}(W; \pi_{g,t}^*)$. Let
Note that $A_{g,t}$ depends on the covariate value $z$, but we suppress its dependence to simplify the exposition. The goal is to construct the uniform confidence band for $\operatorname*{\mathrm{CATT}}_{g,t}(z)$, identified by ${\operatorname*{\mathrm{DR}}}_{g,t}(z) = \operatorname*{\mathbb{E}}[ A_{g,t} \mid Z = z ]$.
For the subsequent discussion, it is convenient to introduce the following notation related to the LQR estimation. For a generic variable $Q$ and a generic integer $\nu \ge 0$, let $\mu_Q^{(\nu)}(z) \coloneqq \operatorname*{\mathbb{E}}^{(\nu)}[Q \mid Z = z]$ denote the $\nu$-th derivative with respect to $z$ of the conditional mean of $Q$ given $Z = z$. As usual, we write $\mu_Q(z) = \mu_Q^{(0)}(z) = \operatorname*{\mathbb{E}}[Q \mid Z = z]$. The LQR estimator of $\mu_Q^{(\nu)}(z)$ for $\nu \in \{ 0, 1, 2 \}$ is defined by
where $\bm{e}_\nu$ is the $3 \times 1$ vector in which the $(\nu + 1)$-th element is $1$ and the rest are $0$, $\bm{r}_2(u) \coloneqq (1, u, u^2)^\top$ is the $3 \times 1$ vector of second order polynomials, $K_Q$ is a kernel function, and $h_Q > 0$ is a bandwidth.
We explain how to obtain the conditional DR estimator $\widehat{\operatorname*{\mathrm{DR}}}_{g,t}(z)$ that appeared in the proposed $(1 - \alpha)$ uniform confidence band for $\operatorname*{\mathrm{CATT}}_{g,t}(z)$ in (ref). Several more technical issues, including the formulas for the standard error and the critical value, are discussed in the following subsections.
We consider a three-step estimation procedure. First, we estimate $\beta_{g,t}^{*}$ and $\pi_{g,t}^*$ via some parametric methods, such as the least squares method and the maximum likelihood estimation. Using the resulting first-stage estimators $\widehat \beta_{g,t}$ and $\widehat \pi_{g,t}$, we compute $\widehat R_{i,g,t} \coloneqq R_{g,t}(W_i; \widehat \pi_{g,t})$ and $\widehat m_{i,g,t} \coloneqq m_{g,t}(X_i; \widehat \beta_{g,t})$ for each $i$. Second, for each $i$, we compute
where $\widehat \mu_G(z)$ is the LQR estimator of $\mu_G(z) = \operatorname*{\mathbb{E}}[ G_g \mid Z = z]$ using a bandwidth $h_G$ and a kernel function $K_G$, and $\widehat \mu_{\widehat R}(z)$ is the LQR estimator of $\mu_R(z) = \operatorname*{\mathbb{E}}[ R_{g,t} \mid Z = z]$ using a bandwidth $h_R$ and a kernel function $K_R$. Specifically, $\widehat \mu_{\widehat R}(z)$ is defined by
and the definition of $\widehat \mu_G(z)$ is analogous. Finally, we obtain the conditional DR estimator $\widehat{\operatorname*{\mathrm{DR}}}_{g,t}(z)$ from the following LQR using a bandwidth $h_A$ and a kernel function $K_A$:
We impose the next assumption on the bandwidths and the kernel functions.
By condition (i), we require a common bandwidth $h$ over the estimation of $\mu_G(z)$, $\mu_R(z)$, and $\mu_A(z)$, which plays an essential role in constructing the critical value. The common bandwidth $h$ should not depend on $(g, t, z)$ for our uniform inference. See Section (ref) and Remark (ref) for further discussion of bandwidth selection.
By condition (ii), we require a common kernel function $K$ over the nonparametric regressions, which is not necessary but greatly facilitates both exposition and theoretical analysis.
We focus on the LQR estimation for the following two reasons. First, the LQR estimation is a standard recommendation for estimating the nonparametric regression function in the kernel smoothing literature due to the boundary adaptive property (fan1996local). Second, it is well known in the literature that, in combination with an appropriate choice of bandwidth, the inference based on the LQR estimation (without analytical bias correction) is numerically identical to the RBC inference based on the bias-corrected LLR estimation. More precisely, the LQR estimator is numerically equivalent to the bias-corrected LLR estimator when the regression function estimation and the bias estimation are carried out with the same appropriate bandwidth (e.g., the IMSE-optimal bandwidth for the LLR estimation), and moreover, the asymptotic variances of the two estimators are identical. This type of RBC inference is simple in the sense that it does not require analytical bias correction, nor does it require adjustment of the standard error due to bias correction, unlike more sophisticated RBC inference. In the literature on RBC inference in kernel smoothing estimation, to the best of our knowledge, only cattaneo2022boundary consider uniform inference based on this simple RBC approach, while other previous studies focus on pointwise inference. Building on their proposal for simple RBC inference, we propose to use the LQR estimation based on the IMSE-optimal bandwidth for the LLR estimation. \phantomsection\Copy{AE-12-1}{ This RBC approach could be generalized to other polynomial orders as in Section 3 of cattaneo2022boundary, but we focus on the LQR-based inference throughout the paper. }
We present an overview of several statistical properties of our estimator, which serve as the bases for the standard error and the critical value. The formal results are relegated to Section (ref).
We will show that the leading term of our estimator is given by
where $\bm{Z} \coloneqq (Z_1, \dots, Z_n)^\top$, $f_Z$ is the density of $Z$, $\Psi_{i,h} \coloneqq ( I_{4,K} - u_{i,h}^2I_{2,K} ) / (I_{4,K} - I_{2,K}^2 ) $ with $u_{i,h} \coloneq (Z_i - z) / h$ and $I_{l,L} \coloneq \int u^l L(u)du$ for a non-negative integer $l$ and a function $L$, and we denote
The second and third terms in $B_{i,g,t}$ originate from the fact that we estimate the nonparametric nuisance parameters $\mu_R(z)$ and $\mu_G(z)$ in the second-stage estimation. Note that the effect of the first-stage estimation does not appear in this (first-order) asymptotic representation because the convergence rates of the first-stage parametric estimators are faster than the nonparametric rate.
Using this asymptotic linear representation, we can derive the asymptotic bias and variance of our estimator. Specifically, we will show that
where
and
with denoting $\sigma_B^2(z) \coloneqq \operatorname*{\mathrm{Var}}[B_{i,g,t} \mid Z_i = z]$ and $\mu_B^{(\nu)}(z) \coloneqq \mu_A^{(\nu)}(z) + [ \mu_E(z) / \mu_R^2(z) ] \mu_R^{(\nu)}(z) - [ \mu_F(z) / \mu_G^2(z) ] \mu_G^{(\nu)}(z)$.
We compute the standard error of our estimator as follows. We start by estimating the density $f_Z(z)$ by some nonparametric method, and let $\widehat f_Z(z)$ denote the resulting estimator. Next, we compute the following variables:
where $\widehat \mu_{\widehat E}(z)$ and $\widehat \mu_{\widehat F}(z)$ denote the LLR estimators of $\mu_E(z)$ and $\mu_F(z)$, respectively. Using these variables, we estimate the conditional variance $\sigma_B^2(z)$ by the following LLR:
where $\widehat U_{i,g,t} \coloneqq \widehat B_{i,g,t} - \widehat \mu_{\widehat B}(Z_i)$ and $\bm{r}_1(u) \coloneqq (1,u)^\top$. Here, the bandwidth and the kernel function can be different from those used for the second- and third-stage estimation. Then, we compute
The asymptotic variance in (ref) can be estimated by $\widehat{\mathcal{V}}_{g,t}(z) / (nh)$, which leads to the following standard error of $\widehat{\operatorname*{\mathrm{DR}}}_{g,t}(z)$: $\widehat{{\operatorname*{\mathrm{SE}}}}_{g,t}(z) \coloneqq \sqrt{ \widehat{\mathcal{V}}_{g,t}(z) / (nh) }$.
We consider two methods for constructing the critical value: (i) an analytical method and (ii) weighted bootstrapping.
\paragraph{Analytical method.}
We will show in Section (ref) that the same type of distributional approximation result as in lee2017doubly holds even in our situation, which is based on the approximation of suprema of empirical processes by suprema of Gaussian processes chernozhukov2014gaussian and several approximation results for suprema of Gaussian processes (piterbarg1996asymptotic; ghosal2000testing). This in turn implies that we can compute the critical value by the same analytical method as in lee2017doubly. \phantomsection\Copy{AE-3-8}{ Specifically, we consider the following critical value for the two-sided symmetric uniform confidence band
where
with $b - a$ corresponding to the length of the interval $\mathcal{I} = [a, b]$. Note that the common bandwidth condition in Assumption (ref)(i) ensures that $\widehat c(1 - \alpha)$ and $a_n^2$ do not depend on $(g, t, z)$. The proposed $(1 - \alpha)$ uniform confidence band over $(g,t,z) \in \mathcal{A}$ is $\widehat{\mathcal{C}} \coloneqq \{ \widehat{\mathcal{C}}_{g,t}(z): (g,t,z) \in \mathcal{A} \}$, where
}
\paragraph{Weighted bootstrapping.}
\phantomsection\Copy{AE-17-1}{ As an alternative to the analytical method, we can consider weighted bootstrap inference. Building on ma2005robust, chen2009efficient, and fan2022estimation, we propose the following algorithm. } For each $b = 1, \dots, B$, we generate a set of IID bootstrap weights $\{V_i^{\star,b}\}_{i=1}^n$ independently of $\{ W_i \}_{i=1}^n$, such that $\operatorname*{\mathbb{E}}[V_i^{\star,b}] = 1$, $\operatorname*{\mathrm{Var}}[V_i^{\star,b}] = 1$, and its distribution has sub-exponential tails. Common choices include a normal random variable with unit mean and unit variance and \phantomsection\Copy{AE-8-1}{ mammen1993bootstrap's (mammen1993bootstrap) wild bootstrap weights such that $\mathbb{P}(V_i^{\star,b} = 2 - c_v) = c_v / \sqrt{5}$ and $\mathbb{P}(V_i^{\star,b} = 1 + c_v) = 1 - c_v / \sqrt{5}$ with $c_v = (\sqrt{5} + 1) / 2$. } \phantomsection\Copy{AE-3-9}{ In each bootstrap repetition, we compute the bootstrapped LQR estimator
and the supremum of the bootstrap counterpart of the studentized statistic
Here, we should use the same $\widehat A_{i,g,t}$, the same bandwidth $h$, and the same kernel function $K$ as the original estimator $\widehat{\operatorname*{\mathrm{DR}}}_{g,t}(z)$. Let $ \widetilde c(1 - \alpha)$ be the empirical $(1 - \alpha)$ quantile of $\{ M^{\star,b} \}_{b = 1}^B$. Note that $\widetilde c(1 - \alpha)$ and $M^{\star,b}$ do not depend on $(g, t, z)$ due to the supremum taken in the definition of $M^{\star,b}$. The $(1 - \alpha)$ uniform confidence band over $(g,t,z) \in \mathcal{A}$ is $\widetilde{\mathcal{C}} \coloneqq \{ \widetilde{\mathcal{C}}_{g,t}(z): (g,t,z) \in \mathcal{A} \}$, where
}
This subsection presents theoretical justifications for the proposed methods.
We impose the following set of mild regularity conditions. Hereafter, to simplify the exposition, we often write “for $(g, t)$” to refer to a generic $(g, t)$ such that $(g, t, z) \in \mathcal{A}$ for $z \in \mathcal{I}$.
Among these assumptions, the undersmoothing condition on the common bandwidth $h$ in Assumption (ref)(iii) is particularly important for our analysis. This assumption ensures that the asymptotic bias $h^4 \mathcal{B}_{g,t}(z)$ is asymptotically negligible when constructing the uniform confidence band. As such, the assumption rules out, for example, computing the LQR estimator $\widehat{\operatorname*{\mathrm{DR}}}_{g,t}(z)$ using the IMSE-optimal bandwidth for the LQR estimation, which is of order $O(n^{-1 /9})$. To fulfill the assumption, we propose using the IMSE-optimal bandwidth for the LLR estimation, not for the LQR estimation, based on the insight of the simple RBC. See Section (ref) for details.
\phantomsection\Copy{AE-21-2}{ Note that our undersmoothing condition accommodates any undersmoothing bandwidth, not limited to our proposal based on the simple RBC approach, as long as the rate required in Assumption (ref)(iii) is satisfied. For example, as in lee2017doubly and fan2022estimation, we can consider a rule-of-thumb adjustment that achieves undersmoothing by shrinking the IMSE-optimal bandwidth for the LQR estimation obtained from the plug-in or cross-validation method by $n^{-\varepsilon}$ for some appropriate $\varepsilon > 0$. However, we prefer the simple RBC approach to the rule-of-thumb adjustment to follow the recent literature that demonstrates the desirable performance of the RBC approach. }
To facilitate our theoretical investigations, Assumption (ref) contains several high-level conditions that can actually be replaced by less restrictive but more complicated conditions. For example, the compact support condition in Assumption (ref)(ii) can be replaced by other conditions that guarantee the existence of technical moments related to the kernel function at the expense of complicated proofs. This in turn implies that commonly used kernel functions (e.g., the Gaussian kernel) should be permissible for use with our analysis.
The next theorem formalizes the asymptotic linear representation and the asymptotic bias and variance formulas described in Section (ref). Hereafter, with an abuse of notation, we often write $o_{\operatorname*{\mathbb{P}}}( \sqrt{ (\log n) / (nh) } )$ to indicate $O_{\operatorname*{\mathbb{P}}}( \sqrt{ (\log n) / (n^{1 + \varepsilon} h) } )$ for some $\varepsilon > 0$ to simplify notation. In addition, \phantomsection\Copy{AE-21-3}{ the theorems presented below treat the bandwidth $h$ as a deterministic sequence, as with many prior studies in the kernel smoothing literature. While investigating the effects of using a data-driven stochastic bandwidth would be interesting, it is beyond the scope of this paper to develop the theory to handle stochastic bandwidths. }
Next, we present a theoretical justification for the uniform confidence band constructed with the analytical method in Section (ref). For this purpose, we consider a uniformly valid distributional approximation result for the studentized statistic. To proceed, we rewrite the standard error as $\widehat{\operatorname*{\mathrm{SE}}}_{g,t}(z) = \widehat{\mathcal{S}}_{g,t}(z) / \sqrt{nh}$, where we denote $\widehat{\mathcal{S}}_{g,t}(z) \coloneqq \sqrt{\widehat{\mathcal{V}}_{g,t}(z)}$ as the estimator of $\mathcal{S}_{g,t} \coloneq \sqrt{\mathcal{V}_{g,t}(z)}$.
We add the following regularity conditions, which is essentially the same as the conditions in Assumption 1 of lee2017doubly. Letting $U_{i,g,t} \coloneqq B_{i,g,t} - \mu_B(Z_i)$ denote the population counterpart of $\widehat U_{i,g,t}$, we write the standard deviation of the $\sqrt{nh}$ times leading term in Theorem (ref) as $\widetilde{\mathcal{S}}_{g,t}(z) \coloneqq h^{-1/2} f_Z^{-1}(z) \sqrt{ \operatorname*{\mathbb{E}} [ \Psi_{i,h}^2 U_{i,g,t}^2 K^2 ( (Z_i - z) / h ) ] }.$ \phantomsection\Copy{AE-20-1}{ This quantity depends on $h$ and also on $n$ through $h$, but we suppress the dependence to simplify notation. }
The next theorem gives a uniformly valid distributional approximation result.
This theorem in turn justifies the use of the analytical critical value $\widehat c(1 - \alpha)$ defined in (ref). To see this, denoting $s_n \coloneqq a_n + s / a_n$, observe that
In words, we can approximate the distribution function of the supremum of the studentized statistic by the right-hand side. Then, it is easy to see that the critical value $\widehat c(1 - \alpha)$ defined in (ref) is the approximated $1 - \alpha$ quantile such that $1 - \alpha = \exp \left( - 2e^{(a_n^2 - \widehat c(1 - \alpha)^2)/2} \right)$. As a result, the $(1 - \alpha)$ uniform confidence band $\widehat{\mathcal{C}}$ defined in (ref) has the desired coverage:
Lastly, we study several theoretical properties for the uniform confidence band constructed with weighted bootstrapping in Section (ref). For this purpose, we make the following assumption on the bootstrap weight. Hereafter, to simplify notation, we suppress the superscript $b$ indicating the bootstrap repetition.
The following theorem gives an asymptotic linear representation for the bootstrap estimator $\widehat{\operatorname*{\mathrm{DR}}}_{g,t}^{\star}(z)$. \phantomsection\Copy{AE-17-3}{ It implies that our weighted bootstrap procedure for CATT is (first-order) asymptotically equivalent to the multiplier bootstrap procedure considered in previous studies (e.g., callaway2021difference). } The proof is almost the same as that of Theorem (ref), and is thus omitted.
Using this result and Corollary 3.1 of chernozhukov2014anti, the next theorem proves the validity of the uniform confidence band constructed with the weighted bootstrap procedure in (ref).
Recall that the construction of our uniform confidence bands relies on the common bandwidth choice and the undersmoothing condition in Assumptions (ref)(i) and (ref)(iii). Here, we propose to use an optimal bandwidth in terms of IMSE for the LLR estimation, instead of that for the LQR estimation. This proposal is based on the insight that the LQR estimation with the IMSE-optimal bandwidth for the LLR estimation can be interpreted as the simple RBC inference, as discussed in Section (ref).
In line with Assumption (ref)(i), we consider the common bandwidth $h_\mathtt{LL} \coloneqq \min_{(g, t)} h_\mathtt{LL}(g, t)$, where $h_\mathtt{LL}(g, t)$ is the IMSE-optimal bandwidth for the LLR estimation for each $(g, t)$, defined as follows. \phantomsection\Copy{AE-23-1}{ Let $\widehat{\operatorname*{\mathrm{DR}}}_{g,t}^{\mathtt{LL}}(z)$ denote the LLR estimator of the DR estimand, whose definition can be found in Appendix (ref). } In Appendix (ref), we show that its asymptotic bias and variance are
Thus, it is easy to see that the (infeasible) IMSE-optimal bandwidth for the LLR estimator is given by
Notice that $h_{\mathtt{LL}}(g,t)$ and thus $h_{\mathtt{LL}}$ are of order $n^{-1/5}$, which satisfies Assumption (ref)(iii).
To construct feasible counterparts of $h_{\mathtt{LL}}$ and $h_{\mathtt{LL}}(g,t)$, we need to estimate the unknown quantities. The estimators for the density $f_Z(z)$ and the conditional variance $\sigma_B^2(z)$ are already available as in Section (ref). The second order derivative $\mu_B^{(2)}(z)$ can be estimated by the $p_B$-th order LPR estimation with some $p_B \ge 2$, where the dependent variable is $\widehat B_{i,g,t}$ and the regressors are the $p_B$-th order polynomials $\bm{r}_{p_B}(Z_i - z)$. Let $\widehat \mu_{\widehat B}^{(2)}(z)$ denote this LPR estimator. Then, we obtain the feasible bandwidths, say $\widehat h_\mathtt{LL} \coloneqq \min_{(g, t)} \widehat h_\mathtt{LL}(g, t)$ and $\widehat h_{\mathtt{LL}}(g,t)$, as the estimated counterparts. Note that using $\widehat h_\mathtt{LL}$ implicitly assumes that all nonparametric functions to be estimated have the same smoothness over all $(g, t, z) \in \mathcal{A}$, which may be restrictive for the same reason as in Remark (ref).
Since our undersmoothing approach is based on the insight of simple RBC, it should be preferable to the traditional rule-of-thumb undersmoothing strategy and the conventional MSE-optimal bandwidth-based inference in terms of achieving both good coverage and a short confidence interval length. In the literature on uniform inference for treatment effect heterogeneity, lee2017doubly and fan2022estimation propose the rule-of-thumb undersmoothing strategy for the LLR estimation, say $\widehat h_{\operatorname*{\mathrm{US}}} \coloneqq \widehat h_{\mathrm{LL}} \cdot n^{1/5} \cdot n^{-2/7}$ in our context. In contrast, several previous studies in the kernel smoothing literature show in theory and numerical analysis that (simple) RBC inference generally leads to better coverage and shorter confidence interval length (e.g., calonico2018effect for pointwise inference in a general setting of kernel smoothing estimation). Thus, we can expect our undersmoothing approach to have the same nice properties. Since our proposal is not based on formal theory but on analogy with previous studies, we thoroughly investigate its performance through a series of Monte Carlo experiments in Appendix (ref), the results of which corroborate our discussion here. We leave more sophisticated RBC inference and bandwidth choices for uniform inference as future topics.
We turn to statistical inference for the aggregated parameter $\theta(z)$ defined in (ref). Since the weighting function $w_{g,t}(z)$ is known or estimable, we can compute the aggregated estimator as follows:
where $\widehat{w}_{g,t}(z) = w_{g,t}(z)$ if $w_{g,t}(z)$ is known, otherwise $\widehat{w}_{g,t}(z)$ is a nonparametric estimator constructed with certain LQR estimation depending on the form of $w_{g,t}(z)$.
To perform the uniform inference for the aggregated parameter $\theta(z)$, we can construct the standard error and the uniform critical value in the same manner as in the case of CATT. To be specific, focusing on the case where $w_{g,t}(z)$ is unknown but estimable, suppose that its estimator satisfies
where $\xi_{i,g,t}$ is an estimable variable whose definition depends on the choice of $w_{g,t}(z)$. In Appendix (ref), we show that this asymptotic linearity holds for a variety of weighting functions of interest. Then, $\widehat{\theta}(z)$ exhibits the same form of asymptotic linearity as in the case of CATT in Theorem (ref):
where $J_i \coloneqq \sum_{g\in\mathcal{G}} \sum_{t=2}^{\mathcal{T}} [ w_{g,t}(z) \cdot B_{i,g,t} + {\operatorname*{\mathrm{DR}}}_{g,t}(z) \cdot \xi_{i,g,t} ]$. Using this result, we can derive the asymptotic bias and variance formulas for the aggregated estimator $\widehat{\theta}(z)$ in the same way as in Theorem (ref). Moreover, the construction of the uniform critical value and the bandwidth selection are essentially the same as in the previous section.
To save space in the main text, we relegate the proof of the asymptotic linear representation in (ref) to Appendix (ref). The appendix also presents the formulas for the standard error and the uniform critical values based on the analytical method and multiplier bootstrapping, and discusses how to implement the uniform inference for concrete summary parameters.
\if00 {
The authors appreciate the valuable comments of the coeditor (Ivan Canay), the associate editor, and two anonymous referees, which greatly improved the paper. The authors also thank Ryo Okui and Toshiki Tsuda for their helpful discussions. This work was supported by JSPS KAKENHI Grant Number 20K13469 and 24KJ1472. } \fi \if10 \fi
The authors report there are no competing interests to declare.